论文解读

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability

语音合成 | 6.6/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6527 字 阅读 →
论文解读

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments

声源定位 | 8.1/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8152 字 阅读 →
论文解读

MedMosaic: A Challenging Large Scale Benchmark of Diverse Medical Audio

音频理解 | 6.4/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8593 字 阅读 →
论文解读

PhoStream: Benchmarking Real-World Streaming for Omnimodal Assistants in Mobile Scenarios

音视频问答 | 7.3/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7522 字 阅读 →
论文解读

Speech-Audio Compositional Attacks on Multimodal LLMs and Their Defense with SALMONN-Guard

音频理解 | 8.3/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6409 字 阅读 →
论文解读

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions

音视频理解 | 9.4/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7829 字 阅读 →
论文解读

Reasoning LLM Improves Speaker Recognition in Long-form TV Dramas

音视频理解 | 7.2/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7157 字 阅读 →
论文解读

Speaker head orientation estimation with a single microphone array using phase spectrogram features

声源定位 | 5.8/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5887 字 阅读 →
论文解读

A Text-Steerable Instrument for Sketching Procedural Soundscapes via Language Models

音乐生成 | 6.8/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5326 字 阅读 →
论文解读

Disentangling Speaker and Language Effects in Cross-Lingual Speaker Verification for Iberian Languages

说话人验证 | 5.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5518 字 阅读 →
论文解读

ORCA: Open-ended Response Correctness Assessment for Audio Question Answering

音频理解 | 7.9/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8254 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-07-02

共分析 16 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 49 分钟 · 24050 字 阅读 →
论文解读

Dilemmadata: On the Interoperability of Heterogeneous Roman Numeral Datasets

数据集 | 10/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4937 字 阅读 →
论文解读

What Counts as an Error? Dual-Reference Benchmarking for Atypical ASR

语音识别 | 7.3/10

 · 更新于 2026-09-25 · 约 25 分钟 · 12422 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-07-01

共分析 35 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 102 分钟 · 50963 字 阅读 →
论文解读

DialogPII: A multilingual dataset of synthetic dialog transcripts to detect personal information

语音识别 | 8.9/10

 · 更新于 2026-09-25 · 约 35 分钟 · 17529 字 阅读 →
论文解读

Underwater Source Detection and Classification for Signal-based Surveillance: Audio Dataset Curation and Cross-Domain Evaluation

数据集 | 7.8/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4494 字 阅读 →
论文解读

Soroll-IA: A Weakly Labeled Audio Dataset for Real-World Industrial Port Monitoring

音频事件检测 | 8.3/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5550 字 阅读 →
论文解读

FoleySet: A Multi-Level Human-Annotated Foley Sound Dataset

音频分类 | 7/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6748 字 阅读 →
论文解读

CN-NewsTTS Bench: a target-level automatic benchmark for raw-input Chinese news TTS pronunciation

语音合成 | 9.2/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8986 字 阅读 →