论文解读

AuRA: Internalizing Audio Understanding into LLMs as LoRA

语音问答 | 7.5/10

 · 更新于 2026-09-06 · 约 7 分钟 · 3100 字 阅读 →
论文解读

DeRA-MOS: Optimizing Text-to-Music Evaluation via Decoupled Listwise Ranking and Modality Alignment

音乐评估 | 8.2/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5161 字 阅读 →
论文解读

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6823 字 阅读 →
论文解读

GlobeAudio: A Multilingual Multicultural Benchmark for Naturalistic Evaluation of Large Audio-Language Models

语音识别 | 7.9/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5486 字 阅读 →
论文解读

KFC-KWS: Keyframe Fusion with CTC for User-Defined Keyword Spotting

关键词检测 | 7.6/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5345 字 阅读 →
论文解读

Optimizing 2D Input Representations and Sub-phase Fusion Strategies for Differential Diagnosis of Asthma and COPD Using CNN- and GRU-Based Networks

数据增强 | 6.8/10

 · 更新于 2026-09-06 · 约 48 分钟 · 23710 字 阅读 →
论文解读

Profy: Interpretable Visualization of Expertise-Dependent Motor Skills Toward Supporting Piano Practice

音乐信息检索 | 6.9/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5822 字 阅读 →
论文解读

RespiraMFM: A Multimodal Foundation Model with Contrastive Audio-Language Alignment for Respiratory Disease Identification

对比学习 | 6.5/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6068 字 阅读 →
论文解读

Spatial-Omni: Spatial Audio Understanding Integration in Multimodal LLMs via FOA Encoding

多模态模型 | 9.4/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5947 字 阅读 →
论文解读

Speech Encoder Fusion for LLM-based Automatic Speech Recognition

语音识别 | 7.2/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5281 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-10

共分析 45 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 126 分钟 · 62690 字 阅读 →
论文解读

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs

语音识别 | 8.8/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6046 字 阅读 →
论文解读

Bridging Traditional Explainability Methods and Multimodal Multilingual Models: An XAI-Based Analysis

语音识别 | 5.4/10

 · 更新于 2026-09-06 · 约 18 分钟 · 8872 字 阅读 →
论文解读

Conan-embedding-v3: Fusing Modality-Specific Models for Omni-Modal Embedding

音频检索 | 7.6/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5579 字 阅读 →
论文解读

Cross-Modal Masking for Robust Silent Speech Synthesis Using sEMG and Lipreading

语音合成 | 7.5/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7716 字 阅读 →
论文解读

Discovering Functionally Selective Brain Regions with a Deep Topographic Multimodal Model

多模态模型 | 8.2/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5464 字 阅读 →
论文解读

HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis

语音合成 | 9/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6999 字 阅读 →
论文解读

Liberating LLM Capabilities in Full-Duplex Speech Models

多模态模型 | 8.7/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8443 字 阅读 →
论文解读

Multi-View Speech Representation Learning for Parkinson's Disease Detection Using Context-guided Cross-modal Attention

自监督学习 | 7.9/10

 · 更新于 2026-09-06 · 约 12 分钟 · 6004 字 阅读 →
论文解读

Paediatric-HGNN: A Hybrid Heterogeneous Graph Neural Network for Detecting Disfluency in Children's Speech via Multiscale Acoustic Fusion

语音合成 | 6.5/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7186 字 阅读 →