语音/音乐/音频论文速递 2026-07-29
语音/音乐/音频论文速递 2026-07-29 共分析 27 篇论文 ⚡ 今日概览 📥 抓取 27 篇 → 🔬 深度分析完成 🏷️ 热门方向 方向 数量 分布 #语音交互 2篇 ██ #语音属性识别 2篇 ██ #语音识别 2篇 ██ #音视频理解 2篇 ██ #音视频问答 2篇 ██ #音频理解 2篇 ██ #Transformer 1篇 █ #声源定位 1篇 █ 📊 论文评分排行榜(27 篇,按分数降序) 排名 论文 总分 分档 文档类型 主任务 🥇 faster-enhancer.c: A Dependency-Free int8 Runtime for S 8.7分 前25% 系统技术报告 #语音增强 🥈 OmniScope: Modality-Decoupled Token Compression for Omn 8.3分 前25% 方法研究 #音视频问答 🥉 AVE-Compass: Towards Holistic Evaluation for Audio-Vide 8.3分 前25% 数据集与基准 #多模态模型 4. Extracting Voice Styles from Frozen TTS Models via Grad 7.9分 前25% 方法研究 #语音克隆 5. CARE: A Multimodal Corpus for Studying Speech and Non-V 7.9分 前25% 数据集与基准 #音视频理解 6. GraphIDyOM: A graph-native Python reimplementation of I 7.8分 前25% 系统技术报告 #音乐理解 7. VAD to the Bone: Ultra-Tiny Speech Activity Detection f 7.4分 前50% 系统技术报告 #语音活动检测 8. Unlocking Spatial Grounding in Large Audio-Visual Retri 7.2分 前50% 方法研究 #声源定位 9. SpeechLLM Meets Federated Learning for End-to-End ASR: 7.1分 前50% 方法研究 #语音识别 10. Joint Text-Audio Alignment for EEG-to-Text Decoding in 7.0分 前50% 方法研究 #语音交互 11. OmniDelta: Skill-Driven Budget Allocation for Token Com 7.0分 前50% 方法研究 #音视频问答 12. Text-Prompted CLAP: Learning Query-Conditioned Audio Re 6.6分 前50% 方法研究 #音频理解 13. Towards Operational Conversational Intelligence: A Spee 6.4分 前50% 系统技术报告 #说话人日志 14. Parallel Decoding Distillation for Fast Image and Video 6.2分 前50% 方法研究 #音视频生成 15. Spacing Out: On the Reliability of Binaural Music Sourc 6.1分 前50% 方法研究 #音乐源分离 16. MyMentorLLM: A psychotherapy GenAI environment with mul 5.9分 前50% 系统技术报告 #语音交互 17. MusiChat: Vibe Composing for Music Creation 5.8分 前50% 系统技术报告 #音乐生成 18. AMRD: Adaptive Multi-Teacher Relational Distillation fo 5.6分 前50% 方法研究 #语音情感识别 19. Multi-Phonation Graph Learning with Self-Supervised Spe 5.5分 前50% 方法研究 #语音属性识别 20. Evaluation of forced alignment of code-mixed speech: th 5.4分 后50% 方法研究 #语音识别 21. A Cross-lingual Comparison of Human and Classification 5.1分 后50% 方法研究 #Transformer 22. From Semantics to Readout: Mechanistic Understanding of 5.1分 后50% 方法研究 #音频理解 23. Finding the noise: Zero-shot AI Music Detection 5.1分 后50% 方法研究 #无监督学习 24. Depression Markers in Speech: An Approach based on Trac 5.1分 后50% 应用研究 #语音属性识别 25. Self-Supervised Audio Representation Learning for Pedia 5.0分 后50% 应用研究 #音频分类 26. DynaBridge: Dynamic Summary-Guided Cross-Task Multimoda 4.9分 后50% 方法研究 #音视频理解 27. Device Invariance using Domain Adaptation on Acoustic S 4.2分 后50% 方法研究 #领域适应 📋 论文列表 🥇 faster-enhancer.c: A Dependency-Free int8 Runtime for Streaming Speech Enhancement on Commodity CPUs 8.7/10 | 创新 1.2/2 | 严谨 1.3/1.5 | 实验 1/1.5 | 清晰 0.9/1 | 影响 0.8/1.5 | 开源 1.5/1.5 | 复现 0.5/0.5 | 工程 1.5/1.5 ...