论文解读

EchoDistill:Alignment Noisy-to-Clean Self-Distillation for Robust Audio LLMs

强化学习 | 9.1/10

 · 更新于 2026-09-09 · 约 13 分钟 · 6317 字 阅读 →
论文解读

MIRAGE: Adaptive Multimodal Gating for Whole-Brain fMRI Encoding

Transformer | 8.2/10

 · 更新于 2026-09-09 · 约 11 分钟 · 5123 字 阅读 →
论文解读

PiAnnotate: A Web Annotation Tool for Piano Fingering, with a Diagnostic Probe

PiAnnotate: A Web Annotation Tool for Piano Fingering, with a Diagnostic Probe

 · 更新于 2026-09-09 · 约 15 分钟 · 7056 字 阅读 →
论文解读

Raon-Speech Technical Report

语音识别 | 6.5/10

 · 更新于 2026-09-09 · 约 11 分钟 · 5464 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-30

共分析 6 篇语音/AI 论文

 · 更新于 2026-09-09 · 约 17 分钟 · 8127 字 阅读 →
论文解读

AgentHijack: Benchmarking Computer Use Agent Robustness to Common Environment Corruptions

AgentHijack: Benchmarking Computer Use Agent Robustness to Common Environment Corruptions

 · 更新于 2026-09-09 · 约 13 分钟 · 6192 字 阅读 →
论文解读

Archon: A Unified Multimodal Model for Holistic Digital Human Generation

语音合成 | 7.5/10

 · 更新于 2026-09-09 · 约 11 分钟 · 5475 字 阅读 →
论文解读

Audio Deepfake Detection with Half-Truth Localisation Using Cross-Attentive Feature Fusion

音频深度伪造检测 | 8.4/10

 · 更新于 2026-09-09 · 约 11 分钟 · 5324 字 阅读 →
论文解读

Audio Jailbreaks in Large Audio-Language Models: Taxonomy, Attack-Defense Analysis, and Cost-Aware Evaluation

多模态模型 | 8.9/10

 · 更新于 2026-09-09 · 约 13 分钟 · 6155 字 阅读 →
论文解读

Benchmarking Single-Factor Physical Video-to-Audio Generation

音频生成 | 9/10

 · 更新于 2026-09-09 · 约 13 分钟 · 6510 字 阅读 →
论文解读

ChildVox: A Speech, Audio, and Large Audio-Language Model Benchmark in Understanding and Characterizing Sound across Childhood

语音识别 | 8/10

 · 更新于 2026-09-09 · 约 11 分钟 · 5167 字 阅读 →
论文解读

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings

音频检索 | 6.5/10

 · 更新于 2026-09-09 · 约 11 分钟 · 5159 字 阅读 →
论文解读

Data-Efficient On-Policy Distillation for Automatic Speech Recognition

语音识别 | 5.1/10

 · 更新于 2026-09-09 · 约 10 分钟 · 4830 字 阅读 →
论文解读

Decoding Strategies for Diffusion-Based ASR: A Systematic Evaluation of Confidence-Based Thresholding

语音识别 | 6.8/10

 · 更新于 2026-09-09 · 约 10 分钟 · 4803 字 阅读 →
论文解读

Dial HEALTHDIAL for Advice: A Multilingual and Multi-Parallel Spoken Dialogue Dataset for Knowledge-Grounded Information Seeking

语音合成 | 8.6/10

 · 更新于 2026-09-09 · 约 12 分钟 · 5895 字 阅读 →
论文解读

DirectorBench: Diagnosing Long-Form Video Generation with Personalized Multi-Agent Evaluation

基准测试 | 9.8/10

 · 更新于 2026-09-09 · 约 14 分钟 · 6831 字 阅读 →
论文解读

HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding

语音合成 | 8.6/10

 · 更新于 2026-09-09 · 约 13 分钟 · 6179 字 阅读 →
论文解读

MELD: Mel-Spectrogram-Based Speech Language Modeling with Discrete Latent Variables

语音合成 | 7.3/10

 · 更新于 2026-09-09 · 约 3 分钟 · 1272 字 阅读 →
论文解读

Mitigating Stethoscope-Induced Shortcuts in Respiratory Sound Classification under Federated Domain Generalization with Causality-Inspired Interventions

音频分类 | 8.5/10

 · 更新于 2026-09-09 · 约 12 分钟 · 5537 字 阅读 →
论文解读

MusTBENCH: Benchmarking and Advancing Temporal Grounding in Music LLMs

音乐生成 | 7.5/10

 · 更新于 2026-09-09 · 约 12 分钟 · 5633 字 阅读 →