论文解读

MMAudioReverbs: Video-Guided Acoustic Modeling for Dereverberation and Room Impulse Response Estimation

语音增强 | 6.0/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5552 字 阅读 →
论文解读

OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs

基准测试 | 8.5/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4370 字 阅读 →
论文解读

OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM

多模态模型 | 7.5/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5484 字 阅读 →
论文解读

Syncphony: Synchronized Audio-to-Video Generation with Diffusion Transformers

音视频 | 8.0/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5201 字 阅读 →
论文解读

TRIBE: TRImodal Brain Encoder for whole-brain fMRI response prediction

脑编码 | 9.5/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4604 字 阅读 →
论文解读

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models

基准测试 | 7.5/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4010 字 阅读 →
论文解读

A cross-species neural foundation model for end-to-end speech decoding

语音识别 | 8.0/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5641 字 阅读 →
论文解读

AUHead: Realistic Emotional Talking Head Generation via Action Units Control

面部动画生成 | 8.0/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5435 字 阅读 →
论文解读

Better Together: Leveraging Unpaired Multimodal Data for Stronger Unimodal Models

多模态模型 | 7.0/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4906 字 阅读 →
论文解读

Closing the Gap Between Text and Speech Understanding in LLMs

语音对话系统 | 7.5/10

 · 更新于 2026-09-24 · 约 17 分钟 · 8119 字 阅读 →
论文解读

FlowBind: Efficient Any-to-Any Generation with Bidirectional Flows

音频生成 | 8.0/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6079 字 阅读 →
论文解读

Learning multimodal dictionary decompositions with group-sparse autoencoders

跨模态 | 7.5/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4031 字 阅读 →
论文解读

MARS-Sep: Multimodal-Aligned Reinforced Sound Separation

语音分离 | 7.5/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5682 字 阅读 →
论文解读

OmniCVR: A Benchmark for Omni-Composed Video Retrieval with Vision, Audio, and Text

音频检索 | 8.5/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5159 字 阅读 →
论文解读

Query-Guided Spatial–Temporal–Frequency Interaction for Music Audio–Visual Question Answering

音频问答 | 7.0/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4714 字 阅读 →
论文解读

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models

基准测试 | 9.0/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4501 字 阅读 →
论文解读

Mapping the Methodological Space of Classroom Interaction Research: Scale, Duration, and Modality in an Age of AI

模型评估 | 6.0/10

 · 更新于 2026-09-24 · 约 6 分钟 · 2848 字 阅读 →
论文解读

Normativity and Productivism: Ableist Intelligence? A Degrowth Analysis of AI Sign Language Translation Tools for Deaf People

语音翻译 | 3.5/10

 · 更新于 2026-09-24 · 约 5 分钟 · 2038 字 阅读 →
论文解读

A Dynamic Gated Cross-Attention Framework for Audio-Text Apparent Personality Analysis

音频分类 | 7.0/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4050 字 阅读 →
论文解读

A LLM-Driven Acoustic Semantic Enriched Framework for Underwater Acoustic Target Recognition

音频分类 | 7.0/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4601 字 阅读 →