论文解读

Speech Quality Embeddings for Improved Detection and Classification of Degradations in Speech Signals

语音质量评估 | 5.8/10

 · 更新于 2026-09-09 · 约 20 分钟 · 9969 字 阅读 →
论文解读

Stage-adaptive Token Selection for Efficient Omni-modal LLMs

多模态模型 | 7.7/10

 · 更新于 2026-09-09 · 约 17 分钟 · 8274 字 阅读 →
论文解读

Synchronization and Turn-Taking in Full-Duplex Speech Dialogue Models

语音对话系统 | 7.6/10

 · 更新于 2026-09-09 · 约 16 分钟 · 7848 字 阅读 →
论文解读

Thinking-while-speaking: A Controlled, Interleaved Reasoning Method for Real-Time Speech Generation

语音对话系统 | 6.9/10

 · 更新于 2026-09-09 · 约 16 分钟 · 8002 字 阅读 →
论文解读

Verifiable Provenance and Watermarking for Generative AI: An Evidentiary Framework for International Operational Law and Domestic Courts

多媒体取证 | 8.6/10

 · 更新于 2026-09-09 · 约 15 分钟 · 7404 字 阅读 →
论文解读

π-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows

长期助手 | 5.2/10

 · 更新于 2026-09-09 · 约 10 分钟 · 4593 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-21

共分析 40 篇语音/AI 论文

 · 更新于 2026-09-09 · 约 144 分钟 · 71894 字 阅读 →
论文解读

A conceptual framework for learning to listen by reward: Curiosity-driven search for novel sources

声源定位 | 5/10

 · 更新于 2026-09-09 · 约 15 分钟 · 7501 字 阅读 →
论文解读

Benchmarking Commercial ASR Systems on Code-Switching Speech: Arabic, Persian, and German

语音识别 | 6.2/10

 · 更新于 2026-09-09 · 约 19 分钟 · 9201 字 阅读 →
论文解读

Can Large Language Models Reliably Correct Errors in Low-Resource ASR? A Contamination-Aware Case Study on West Frisian

语音识别 | 7.5/10

 · 更新于 2026-09-09 · 约 19 分钟 · 9491 字 阅读 →
论文解读

CounterFlow: A Two-Phase Inference-Time Sampling for Counterfactual Video Foley Generation

音频生成 | 6/10

 · 更新于 2026-09-09 · 约 16 分钟 · 7631 字 阅读 →
论文解读

Cross-Talk Speech Reduction, by Separation, for Separation

语音分离 | 8.3/10

 · 更新于 2026-09-09 · 约 20 分钟 · 9536 字 阅读 →
论文解读

DASM: Domain-Aware Sharpness Minimization for Multi-Domain Voice Stream Steganalysis

语音伪造检测 | 7/10

 · 更新于 2026-09-09 · 约 15 分钟 · 7015 字 阅读 →
论文解读

EMO-BOOST: Emotion-Augmented Audio-Visual Features for Improved Generalization in Deepfake Detection

音频深度伪造检测 | 7.2/10

 · 更新于 2026-09-09 · 约 17 分钟 · 8115 字 阅读 →
论文解读

Executable Boundary Contracts for Sound Event Traces

音频事件检测 | 8.4/10

 · 更新于 2026-09-09 · 约 19 分钟 · 9164 字 阅读 →
论文解读

Fast Multichannel NMF with Block-Diagonal Spatial Covariance Matrices for Efficient Blind Source Separation Using Distributed Microphone Arrays

语音分离 | 6.5/10

 · 更新于 2026-09-09 · 约 14 分钟 · 6659 字 阅读 →
论文解读

FormalASR: End-to-End Spoken Chinese to Formal Text

语音识别 | 6/10

 · 更新于 2026-09-09 · 约 15 分钟 · 7432 字 阅读 →
论文解读

GroupAffect-4: A Multimodal Dataset of Four-Person Collaborative Interaction

数据集 | 6.8/10

 · 更新于 2026-09-09 · 约 19 分钟 · 9307 字 阅读 →
论文解读

Heterogeneity-Aware Dataset Scheduling for Efficient Audio Large Language Model Training

音频问答 | 7/10

 · 更新于 2026-09-09 · 约 17 分钟 · 8063 字 阅读 →
论文解读

Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation

语音识别 | 6.8/10

 · 更新于 2026-09-09 · 约 22 分钟 · 10575 字 阅读 →