每日自动抓取 arxiv/huggingface 最新语音/音乐/音频AI 论文,AI 深度分析后发布
语音/音乐/音频论文速递 2026-09-04
语音/音乐/音频论文速递 2026-09-04 共分析 30 篇论文 ⚡ 今日概览 ✅ 筛选入选 30 篇 → 🔬 深度分析完成 🏷️ 热门方向 方向 数量 分布 #语音识别 6 篇 ██████ #语音增强 5 篇 █████ #语音合成 3 篇 ███ #语音交互 2 篇 ██ #语音质量评估 2 篇 ██ #声源定位 1 篇 █ #语音情感识别 1 篇 █ #语音编码 1 篇 █ 📊 论文评分排行榜(30 篇,按分数降序) 排名 论文 总分 分档 文档类型 主任务 🥇 ToolDF: Tool-Integrated Reasoning for Mixed… 8.4 前25% 方法研究 #音频伪造检测 🥈 Geometric Ceilings on Time-Frequency Masking for… 7.9 前25% 理论研究 #音乐源分离 🥉 CRAW: Codec Robust Audio Watermarking 7.7 前25% 方法研究 #音频水印 4. The 2026 PNPL Competition: Word Classification and… 7.5 前25% 数据集与基准 #语音识别 5. Test-time adaptation for speech enhancement with an… 7.5 前25% 方法研究 #语音增强 6. Alignment-Free Text-Audiobox for Voice Dubbing and… 7.5 前25% 系统技术报告 #语音合成 7. Listen to the Latents: Self-Correcting Speech… 7.2 前50% 方法研究 #语音识别 8. StreamWSR: Streamable and Lightweight Waveform-Domain… 7.2 前50% 方法研究 #语音增强 9. DuplexSpeechBench-IFEval: Evaluating Implicit… 7.2 前50% 数据集与基准 #语音交互 10. Building and Evaluating Fixed-Voice Thai TTS from… 7.2 前50% 系统技术报告 #语音合成 11. PACodec: A Low-bitrate Neural Speech Codec with… 7.1 前50% 方法研究 #语音编码 12. Decoupling Turn-Taking from Semantics: A Decoupled… 7.0 前50% 方法研究 #语音交互 13. Summary of the ChinaVoices Challenge 2026: Data, Tasks… 6.9 前50% 数据集与基准 #语音识别 14. The Attention Triangle in Audio-Video Models 6.9 前50% 方法研究 #音视频生成 15. Compressing Streaming Neural Audio Encoders via Latent… 6.9 前50% 方法研究 #语音识别 16. Dual-Form ASR: Semantics-Aware Inverse Text… 6.8 前50% 方法研究 #语音识别 17. Deep Neural Compression for RIR-Characterized Acoustic… 6.6 前50% 方法研究 #音频编码 18. SISER: Speaker-Invariant Speech Emotion Recognition… 6.5 前50% 方法研究 #语音情感识别 19. Masked Autoregressive Speech Enhancement with… 6.4 前50% 方法研究 #语音增强 20. Last Translation Benchmark 6.4 前50% 数据集与基准 #语音翻译 21. Neural Music Enhancement with Dual Time-Frequency… 6.2 前50% 方法研究 #语音增强 22. CAPQ-FAST: Content-Adaptive Perceived Quality… 6.2 前50% 应用研究 #音频质量评估 23. Beyond .WAV: Design and Software Verification of… 6.1 前50% 系统技术报告 #语音质量评估 24. Broadband Acoustic Intensity Direction Estimation with… 6.1 前50% 方法研究 #声源定位 25. Local Chord Corruption Is Not Recognizer Replay: Chord… 6.1 前50% 方法研究 #音乐生成 26. VoxReason: Listener-Free Evaluation of Source-Grounded… 5.9 前50% 数据集与基准 #语音合成 27. Is Semantics Enough for Speech Mean Opinion Score… 5.9 前50% 方法研究 #语音质量评估 28. Fairness Evaluation of Edge-AI Implementation for… 5.6 前50% 应用研究 #语音识别 29. StrixAE: An Intelligent Agent for Audio Enhancement… 5.5 前50% 系统技术报告 #语音增强 30. Opening mind by opening architecture: analysis… 4.5 后 50% 系统技术报告 #音频生成 📋 论文列表 🥇 真假混在一起时,只给整段打分为什么不够:ToolDF 的分诊式推理 英文题目:ToolDF: Tool-Integrated Reasoning for Mixed-Authenticity Audio Deepfake Detection ...
语音/音乐/音频论文速递 2026-09-03
语音/音乐/音频论文速递 2026-09-03 共分析 18 篇论文 ⚡ 今日概览 ✅ 筛选入选 18 篇 → 🔬 深度分析完成 🏷️ 热门方向 方向 数量 分布 #语音识别 6 篇 ██████ #语音合成 2 篇 ██ #音频分类 2 篇 ██ #声源定位 1 篇 █ #语音增强 1 篇 █ #语音情感识别 1 篇 █ #音乐生成 1 篇 █ #音视频理解 1 篇 █ 📊 论文评分排行榜(18 篇,按分数降序) 排名 论文 总分 分档 文档类型 主任务 🥇 AVERT: Audio-Verified Adjudication for Spoken Dialogue… 7.6 前25% 方法研究 #语音识别 🥈 Hearing the Whispers: Black-Box Membership Inference… 7.5 前25% 方法研究 #语音合成 🥉 The Missing Temporal Link: Temporal Context Routing… 7.5 前25% 方法研究 #音视频生成 4. VibeVoice-ASR-Streaming Technical Report 7.5 前25% 系统技术报告 #语音识别 5. A Common Measure of Communication for Speech Brain… 7.4 前50% 方法研究 #语音识别 6. SonicCaps: Large-Scale Diverse and Fine-Grained… 7.2 前50% 数据集与基准 #音频检索 7. Understanding Automatic Mixing: A Subtask-Oriented… 7.0 前50% 应用研究 #音乐生成 8. Scalable Direction-Following TTS via Voice Impression… 6.8 前50% 方法研究 #语音合成 9. Efficient Passive Acoustic Monitoring of Killer Whales… 6.7 前50% 应用研究 #音频分类 10. Auditory Illusion Benchmark for Large Audio Language… 6.4 前50% 数据集与基准 #音频理解 11. ARFT: A Synchronized Multimodal RF-Acoustic Dataset… 6.4 前50% 数据集与基准 #声源定位 12. Sensing Bone-Conducted Speech with Earbuds 6.3 前50% 应用研究 #语音增强 13. PhoenixNest-Video: Evidence-Grounded Multimodal Agent… 6.3 前50% 方法研究 #音视频理解 14. SpeakPay: Domain-Adaptive LoRA Fine-Tuning of Whisper… 6.1 前50% 应用研究 #语音识别 15. Removing Speech, Keeping Activities: A Privacy… 5.9 前50% 应用研究 #音频分类 16. Choosing a PEFT Variant for Per-Patient Dysarthric ASR… 5.9 前50% 应用研究 #语音识别 17. VAANI Noise Event Dataset: A curated spontaneous… 5.8 前50% 数据集与基准 #语音识别 18. Predictors of Loneliness in Older Adults Using… 4.9 后 50% 应用研究 #语音情感识别 📋 论文列表 🥇 别让音频去写作,让它来验货:AVERT 对口语状态跟踪的裁决式修正 英文题目:AVERT: Audio-Verified Adjudication for Spoken Dialogue State Tracking ...
语音/音乐/音频论文速递 2026-09-02
语音/音乐/音频论文速递 2026-09-02 共分析 23 篇论文 ⚡ 今日概览 ✅ 筛选入选 23 篇 → 🔬 深度分析完成 🏷️ 热门方向 方向 数量 分布 #音频分类 3 篇 ███ #多模态模型 2 篇 ██ #语音合成 2 篇 ██ #语音增强 2 篇 ██ #语音情感识别 2 篇 ██ #音频分离 2 篇 ██ #语音交互 1 篇 █ #语音伪造检测 1 篇 █ 📊 论文评分排行榜(23 篇,按分数降序) 排名 论文 总分 分档 文档类型 主任务 🥇 BiMTokenizer: Preserving Semantic-Acoustic Balance in… 8.3 前25% 方法研究 #语音编码 🥈 A Composable Evaluation System for Reproducible Omni… 8.3 前25% 系统技术报告 #多模态模型 🥉 Phrase-Localized Language-Contrastive Guidance… 8.2 前25% 方法研究 #语音合成 4. Zero-Shot Respiratory Sound Classification through LLM… 7.8 前25% 方法研究 #音频分类 5. A Unified Uncertainty-Aware Back-End for Speaker… 7.8 前25% 方法研究 #说话人验证 6. ABSE-NET: A Lightweight Neural Model for Active… 7.5 前25% 方法研究 #语音增强 7. TUTTI: Toward generalizable audio-to-score… 7.4 前50% 方法研究 #音乐转录 8. Perceptible or Not? Diagnosing Passive Fingerprints… 7.4 前50% 方法研究 #语音伪造检测 9. Ready to Speak: Aligning LLMs for TTS-Friendly Text… 7.4 前50% 方法研究 #语音合成 10. On the Human and Computer Alignment of Attribute-Based… 7.3 前50% 数据集与基准 #音乐检索 11. TimeSteer: Inference-Time Speech Scheduling in Joint… 7.2 前50% 方法研究 #音视频生成 12. VoiceLongMemEval: Do Assistants Remember How You… 7.1 前50% 数据集与基准 #语音情感识别 13. Cleaner Speech, Weaker Generalization: Revisiting Pitt… 7.0 前50% 数据集与基准 #音频分类 14. TAG-Bench: Benchmarking Temporal Audio Grounding in… 6.9 前50% 数据集与基准 #音频理解 15. Artificial Rosetta Stone: Constrained Maximum A… 6.8 前50% 方法研究 #音乐理解 16. XVAE-WMT: Explainable Wavelet-Temporal Variational… 6.6 前50% 方法研究 #音频分离 17. U-PAST: A Phase-Aware Audio Spectrogram Transformer-U… 6.4 前50% 方法研究 #语音增强 18. Conversation Coach: A Voice-enabled AI System that… 6.3 前50% 系统技术报告 #语音交互 19. Soft Posterior Speaker Injection for Multi-Talker… 6.2 前50% 方法研究 #语音识别 20. Heard but Not Heeded: Paralinguistic Information… 6.0 前50% 方法研究 #语音情感识别 21. Ontology-based Target Sound Extraction 5.7 前50% 方法研究 #音频分离 22. MADS: A Multiview Acoustic Descriptor Set Beyond… 5.7 前50% 方法研究 #音频分类 23. TEIDAN: A Multilingual Multiparty Dialogue Corpus 5.1 后 50% 数据集与基准 #多模态模型 📋 论文列表 🥇 单塔为何还能赢双塔:BiMTokenizer 用双向状态与固定格点重做 1.1 kbps 的语义-声学平衡 英文题目:BiMTokenizer: Preserving Semantic-Acoustic Balance in Low-Bitrate Speech Tokenization via Bidirectional State-Space Modeling ...
语音/音乐/音频论文速递 2026-09-01
语音/音乐/音频论文速递 2026-09-01 共分析 49 篇论文 ⚡ 今日概览 ✅ 筛选入选 49 篇 → 🔬 深度分析完成 🏷️ 热门方向 方向 数量 分布 #语音识别 8 篇 ████████ #语音情感识别 5 篇 █████ #音乐生成 5 篇 █████ #语音合成 4 篇 ████ #语音增强 4 篇 ████ #音乐理解 4 篇 ████ #说话人日志 2 篇 ██ #音乐转录 2 篇 ██ 📊 论文评分排行榜(49 篇,按分数降序) 排名 论文 总分 分档 文档类型 主任务 🥇 Anchoring Speech with Semantics: A Multimodal Adapter… 8.4 前25% 方法研究 #语音识别 🥈 Playability-Aware Audio-to-Tablature Guitar… 8.2 前25% 方法研究 #音乐转录 🥉 When Vocal Tone and Literal Meaning Diverge: An… 8.1 前25% 数据集与基准 #语音情感识别 4. Self-Supervised Pretext Tasks for Infant Cry Analysis… 8.1 前25% 方法研究 #音频分类 5. ImageEval 2026: Culturally Grounded Arabic Multimodal… 8.1 前25% 数据集与基准 #语音识别 6. Perceptually Better, Semantically Worse: Measuring… 8.0 前25% 数据集与基准 #语音增强 7. Vocal Music under Phoneme-Conditional Analysis 8.0 前25% 方法研究 #音乐理解 8. Likelihood-Constrained Acoustic Reranking for Training… 7.8 前25% 方法研究 #语音识别 9. Using Prosody to Predict Syntactic Structure 7.6 前25% 方法研究 #Transformer 10. When Does Predictor-Based RL Align with Human… 7.6 前25% 方法研究 #语音合成 11. V2TATC: A Joint Voice-Trajectory Embedding Framework… 7.5 前25% 方法研究 #对比学习 12. Sequential Trajectories and Simultaneous Blending… 7.5 前25% 方法研究 #语音合成 13. No Detectable Change in Side-Level WER from Prompt… 7.4 前50% 应用研究 #语音识别 14. VocalAffectBench: Evaluating Vocal Emotion Recognition… 7.4 前50% 数据集与基准 #语音情感识别 15. When Models Hear What They Expect: Diagnosing Prosodic… 7.4 前50% 方法研究 #语音情感识别 16. Context-Aware Interleaved Batching for WhisperX 7.4 前50% 方法研究 #语音识别 17. Enabling Proactive Spoken Turns via a Generalized… 7.3 前50% 方法研究 #语音交互 18. MusGU+: Toward a Musician-Centered Evaluation… 7.3 前50% 数据集与基准 #音乐生成 19. Stride-k Subsampling: Train-Free Audio Token Reduction… 7.2 前50% 方法研究 #语音识别 20. TEMPO: Temporally-grounded Multi-task Post-training… 7.1 前50% 方法研究 #音频事件检测 21. VIBE: Video Instruction-aligned Background music… 7.1 前50% 方法研究 #音乐生成 22. Diagnose, Then Refine: A Closed-Loop TTS System with… 7.0 前50% 方法研究 #语音合成 23. HEAR Who Said What: Unlocking Speaker-Attributed… 7.0 前50% 数据集与基准 #说话人日志 24. Evidence-Bounded Mental Health Reasoning from… 7.0 前50% 数据集与基准 #语音情感识别 25. SPHERE: Automatic Music Upmixing via Audio Language… 6.9 前50% 方法研究 #音乐生成 26. Linguistic Distance Segregates Latent Representations… 6.9 前50% 应用研究 #语音识别 27. Closing the Verification Loop: Self-Check Captioning… 6.8 前50% 方法研究 #音频字幕生成 28. Conjoint Audio-to-Spikes Encoding and Processing for… 6.8 前50% 方法研究 #语音识别 29. Neural Multichannel Distant Speaker Diarization and… 6.6 前50% 方法研究 #说话人日志 30. PhysWave: Physics-Guided Latent Diffusion Models for… 6.6 前50% 方法研究 #音频生成 31. What Are You Listening to? Temporal Music Grounding… 6.5 前50% 数据集与基准 #音乐理解 32. CoJEPA: Combining Contrastive Learning and JEPA for… 6.5 前50% 方法研究 #音乐理解 33. Ouroboros: Self-Referential Backdoor Attacks on Speech… 6.4 前50% 方法研究 #语音增强 34. Textual Acoustic Grounding for Generalizable LLM-Based… 6.4 前50% 方法研究 #语音伪造检测 35. Towards Balanced Spectral Reconstruction: Spectrally… 6.4 前50% 方法研究 #语音增强 36. Accurate Plate Reverb Parameter Estimation Using Two… 6.3 前50% 方法研究 #音乐理解 37. Beyond Speech: Dual-Domain SSL Fusion for Unified All… 6.3 前50% 系统技术报告 #音频伪造检测 38. Weakly Supervised Tabla Stroke Transcription via an… 6.3 前50% 方法研究 #音乐转录 39. Language-Statistical Analysis of Neural Audio Codec… 6.3 前50% 方法研究 #语音编码 40. Parallel Time-Band Mixing with Learned Observation… 6.2 前50% 方法研究 #语音增强 41. How Well Do Generative Music Models Follow Emotion… 6.0 前50% 数据集与基准 #音乐生成 42. Multimodal Adaptive Expert Selection with Text Routing… 6.0 前50% 方法研究 #音视频理解 43. Evoking Harmony via Convolution 5.9 前50% 方法研究 #音乐生成 44. DreamX-Creator: Democratizing Native Audio-Video… 5.9 前50% 模型报告 #扩散模型 45. Leveraging Bayesian Optimization for Array Shape Self… 5.7 前50% 方法研究 #声源定位 46. Opinionated, Hesitant and Stressed: Three Studies of… 5.6 前50% 应用研究 #语音情感识别 47. Disentangling Representation using Attributes-based… 5.5 前50% 方法研究 #音频分类 48. Decoupled Latent Flow Matching for Few-Step Joint… 5.3 后 50% 方法研究 #音乐源分离 49. Audio-Driven Adversarial Defense for 3D Talking Face… 4.3 后 50% 方法研究 #语音合成 📋 论文列表 🥇 当解码器缺的不是声音,而是下一词的语义证据 英文题目:Anchoring Speech with Semantics: A Multimodal Adapter Mechanism for Automatic Speech Recognition in Low-Resource Languages ...
语音/音乐/音频论文速递 2026-08-31
语音/音乐/音频论文速递 2026-08-31 共分析 22 篇论文 ⚡ 今日概览 ✅ 筛选入选 22 篇 → 🔬 深度分析完成 🏷️ 热门方向 方向 数量 分布 #音乐理解 3 篇 ███ #语音交互 2 篇 ██ #语音增强 2 篇 ██ #语音识别 2 篇 ██ #音视频理解 2 篇 ██ #音频分类 2 篇 ██ #音频生成 2 篇 ██ #语音情感识别 1 篇 █ 📊 论文评分排行榜(22 篇,按分数降序) 排名 论文 总分 分档 文档类型 主任务 🥇 Phoneme- and Word-Level Metrics Using Self-Supervised… 9.1 前10% 方法研究 #语音识别 🥈 Not all generalisation failures can be bought back… 8.5 前25% 方法研究 #音频理解 🥉 SURE-Challenge: Evaluating Speech Evidence Before… 8.3 前25% 数据集与基准 #语音识别 4. Exploring the Design Space of Representation Learning… 8.0 前25% 方法研究 #音频检索 5. Alias-Free Oscillator Synchronization via Additive… 7.9 前25% 系统技术报告 #音乐生成 6. PolyMap: A 64-Channel Polyphonic Guitar Pickup System 7.7 前25% 系统技术报告 #音乐源分离 7. Multirate State Space Models for End-to-End Processing… 7.5 前25% 方法研究 #语音增强 8. A Mixed-Behavior Vote Model for Multimedia Subjective… 7.4 前50% 方法研究 #语音质量评估 9. A Frequency-Domain Artificial Reverberator Plug-In 7.2 前50% 系统技术报告 #音频生成 10. Low-Power End-to-End Cochlear Implant Speech Denoising… 6.7 前50% 方法研究 #语音增强 11. Compositional Failure in Audio-Visual LLMs: Late-Layer… 6.5 前50% 方法研究 #音视频理解 12. Evaluating Loss Functions in Differentiable Out-of… 6.4 前50% 方法研究 #音频生成 13. A Shaky Voice Is Not Always a Dodge: Benchmarking… 6.2 前50% 数据集与基准 #语音情感识别 14. MuSP-Bench: Advanced Multimodal Benchmarking of Music… 6.2 前50% 数据集与基准 #音乐理解 15. Effects of HRTF Augmentation on Predicted Spatial… 6.1 前50% 方法研究 #音乐理解 16. Auditing Generative Audio Calls for Known-Task Audio… 6.0 前50% 方法研究 #音频分类 17. Is Prosody Lost in Translation? Fine-Grained Cross… 6.0 前50% 数据集与基准 #语音翻译 18. Predicting Turn-Taking Outcomes in Multi-Party… 6.0 前50% 方法研究 #语音交互 19. SETU: An Agentic Ecosystem for Multilingual, Persona… 5.5 前50% 系统技术报告 #音视频理解 20. Klangfarbenakkord and Klangfarbenharmonien Metric… 5.5 前50% 理论研究 #音乐理解 21. Effectiveness of IoT and Deep Learning for Detection… 5.1 后 50% 应用研究 #音频分类 22. When Robots Mishear Us: Mapping the Safety Risks of… 5.1 后 50% 应用研究 #语音交互 📋 论文列表 🥇 没有金标准时,怎么判断对齐切得准不准 英文题目:Phoneme- and Word-Level Metrics Using Self-Supervised Speech Representations for Forced Alignment Evaluation ...
语音/音乐/音频论文速递 2026-08-29
语音/音乐/音频论文速递 2026-08-29 共分析 29 篇论文 ⚡ 今日概览 ✅ 筛选入选 29 篇 → 🔬 深度分析完成 🏷️ 热门方向 方向 数量 分布 #语音识别 5 篇 █████ #语音交互 3 篇 ███ #音视频理解 3 篇 ███ #音频检索 2 篇 ██ #音频理解 2 篇 ██ #LoRA 1 篇 █ #多模态模型 1 篇 █ #空间音频 1 篇 █ 📊 论文评分排行榜(29 篇,按分数降序) 排名 论文 总分 分档 文档类型 主任务 🥇 AudioSpan: Spanning the Duration and Depth of Audio… 8.7 前25% 数据集与基准 #音频理解 🥈 StreamAV-Bench: A Comprehensive Benchmark for… 8.2 前25% 数据集与基准 #音频生成 🥉 AfriSwitch: A Benchmark for In-the-Wild African Code… 8.1 前25% 数据集与基准 #语音识别 4. Said Aloud, Read Different: Cross-Modal Instability in… 8.1 前25% 数据集与基准 #音视频问答 5. Your Voice Cloning System is Secretly a Voice… 8.0 前25% 方法研究 #语音转换 6. Vagdhenu: A Vrutta (Meter) Aware Shloka-to-Chant (TTS)… 7.9 前25% 系统技术报告 #语音合成 7. Modality Maturity Index: A benchmark for assessing… 7.8 前25% 数据集与基准 #音视频理解 8. Emotion Understanding in Streaming Video with… 7.7 前25% 方法研究 #音视频理解 9. SpeechGym: An Audio-Native Gym for Training Voice… 7.6 前25% 系统技术报告 #语音交互 10. Omni-Interactive Universal Embedder 7.5 前25% 方法研究 #音频检索 11. Multi2AV-Safety: Benchmarking Safety in Multimodal-to… 7.3 前50% 数据集与基准 #音视频生成 12. Mapping Written Words to Spoken Words in a Different… 7.3 前50% 方法研究 #音频检索 13. Towards Interpretable Depression Detection: Linking… 7.1 前50% 系统技术报告 #语音情感识别 14. Letters hide the truth from our eyes: English… 6.8 前50% 方法研究 #语音识别 15. Interpretable, Fairly Evaluated Automated L2 Speaking… 6.7 前50% 方法研究 #语音质量评估 16. When Text Misleads: Inconsistent-Aware Reasoning for… 6.6 前50% 数据集与基准 #语音交互 17. From Sound to Symptom: Real-Time Respiratory Signal… 6.4 前50% 系统技术报告 #音频事件检测 18. Scaling phoneme-based TTS augmentation for ASR: A… 6.2 前50% 方法研究 #语音识别 19. Direct or Mediated? Task-Dependent Audio Information… 6.1 前50% 方法研究 #音频理解 20. Attention-Guided Reliability Scaling for Contrastive… 6.0 前50% 方法研究 #语音识别 21. Soft Active Electromyography Interface for Machine… 5.9 前50% 系统技术报告 #语音识别 22. A Reranker for Orchestrating Heterogeneous Speech and… 5.8 前50% 方法研究 #LoRA 23. Decay-Region Group Delay as a Forensic Cue for AI… 5.8 前50% 方法研究 #音频伪造检测 24. Recovering Expert Critic-Sourced Network Adjacency… 5.8 前50% 方法研究 #音乐推荐 25. Benchmarking_Fast_Domain_Adaptation_for_Unsupervised_S… 5.5 前50% 数据集与基准 #语音编码 26. GAN-based Joint Dereverberation and Directional… 5.3 后 50% 方法研究 #空间音频 27. Real-TurnTurk: A Multimodal Turkish Corpus for Turn… 5.1 后 50% 数据集与基准 #多模态模型 28. A Safety-Gated Multimodal AI Backend for Mental-Health… 5.0 后 50% 系统技术报告 #语音交互 29. How AI Experiences Art: Emergent Aesthetic Structure… 4.0 后 50% 方法研究 #音视频理解 📋 论文列表 🥇 长音频不是更长的短音频:AudioSpan 如何逼模型证明它真的听见了 英文题目:AudioSpan: Spanning the Duration and Depth of Audio Comprehension ...
语音/音乐/音频论文速递 2026-08-27
语音/音乐/音频论文速递 2026-08-27 共分析 21 篇论文 ⚡ 今日概览 ✅ 筛选入选 21 篇 → 🔬 深度分析完成 🏷️ 热门方向 方向 数量 分布 #语音识别 4 篇 ████ #回声消除 2 篇 ██ #语音交互 2 篇 ██ #音频分类 2 篇 ██ #音频理解 2 篇 ██ #基准测试 1 篇 █ #空间音频 1 篇 █ #语音伪造检测 1 篇 █ 📊 论文评分排行榜(21 篇,按分数降序) 排名 论文 总分 分档 文档类型 主任务 🥇 AllMusicCaps: Album Reviews as Complementary… 9.1 前10% 方法研究 #音乐检索 🥈 LibriBrain100: One Hundred Hours of Broad and Deep MEG… 8.9 前25% 数据集与基准 #基准测试 🥉 What Do Audio-Visual Synchronization Metrics Actually… 8.8 前25% 方法研究 #音视频理解 4. Why ML-based cough models do not generalize: a… 8.8 前25% 应用研究 #音频分类 5. Lost but not erased: Finding traces of a forgotten… 8.6 前25% 方法研究 #语音识别 6. TurnBench: A Multi-Domain Benchmark for Turn-Taking… 8.5 前25% 数据集与基准 #语音交互 7. Knowledge Distillation for Efficient Acoustic Echo… 8.5 前25% 方法研究 #回声消除 8. Domain-Adaptive ASR for Telephony AI Agents: Fine… 7.9 前25% 系统技术报告 #语音识别 9. CSAVocoder: A Causal Spatial Audio Vocoder Towards… 7.6 前25% 系统技术报告 #空间音频 10. SPECTRA: Subspace-Preserving Embedding Calibration… 7.2 前50% 方法研究 #音频分类 11. Dissonance Spectrum explicitly models perceptual… 7.2 前50% 方法研究 #音乐理解 12. AudioLens: Multi-Perspective Speech Clustering with… 7.1 前50% 方法研究 #音频理解 13. Super Star: Towards Streaming Real-time Interactive… 6.9 前50% 系统技术报告 #音视频交互 14. Can We Read the Mind of an Audio LLM? A Verbalizable… 6.6 前50% 方法研究 #音频理解 15. VoiceMem: Streaming Dual-Brain Memory for Real-Time… 6.6 前50% 系统技术报告 #语音交互 16. A Training-Free Proactive Defense Against Partial… 6.5 前50% 方法研究 #语音伪造检测 17. Combining Self-Embedding Audio Watermarking with Ultra… 6.4 前50% 方法研究 #音频水印 18. Generative vs. Encoder Large Language Models for ASR… 6.1 前50% 应用研究 #语音质量评估 19. Mandarin Humorous Homophone Recognition and… 5.7 前50% 方法研究 #语音识别 20. Acoustic Echo Control Based on Sound Object… 4.9 后50% 方法研究 #回声消除 21. Fine-Tuning Whisper for Automatic Speech Recognition… 4.7 后50% 应用研究 #语音识别 📋 论文列表 🥇 把乐评变成检索监督,关键不在多而在语域对位 英文题目:AllMusicCaps: Album Reviews as Complementary Supervision for Music CLAP ...
语音/音乐/音频论文速递 2026-08-26
语音/音乐/音频论文速递 2026-08-26 共分析 26 篇论文 ⚡ 今日概览 ✅ 筛选入选 26 篇 → 🔬 深度分析完成 🏷️ 热门方向 方向 数量 分布 #音频理解 7 篇 ███████ #空间音频 3 篇 ███ #语音交互 2 篇 ██ #语音合成 2 篇 ██ #音乐生成 2 篇 ██ #音视频理解 2 篇 ██ #声源定位 1 篇 █ #语音伪造检测 1 篇 █ 📊 论文评分排行榜(26 篇,按分数降序) 排名 论文 总分 分档 文档类型 主任务 🥇 EmoTra-TTS: Smooth Intra-Utterance Emotion Transitions… 9.6 前10% 方法研究 #语音合成 🥈 LAION-BVD: A 10-Million-Hour Open Video Dataset for… 9.4 前10% 数据集与基准 #音视频理解 🥉 The ISCSLP 2026 Real-World Audio-Visual Speech… 8.8 前25% 数据集与基准 #音视频语音分离 4. EM-KalmanNet: Learned Expectation-Maximization for… 8.8 前25% 方法研究 #声源定位 5. EXAM\(^2\): \(\underline{Ex}tending\) \(\underline{A}udio\)… 8.7 前25% 数据集与基准 #音频理解 6. FireRedAudio: A General-Purpose Audio Language Model… 8.7 前25% 模型报告 #音频理解 7. Arbitrary Polygon Oscillator: Generalizing Polygonal… 8.7 前25% 方法研究 #音乐生成 8. Relative Time Intervals Representation for Word-level… 8.6 前25% 方法研究 #语音识别 9. REDnet: Recursive Encoder and Decoder for Speech… 8.2 前25% 方法研究 #语音分离 10. From local kernels to global form: modeling the… 8.2 前25% 理论研究 #音乐理解 11. Lost in Speech: Trilingual Spoken Hallucination… 8.2 前25% 数据集与基准 #音频理解 12. Weakly Supervised Seafloor Segmentation for Seagrass… 8.1 前25% 方法研究 #音频理解 13. SonarLLM: A Native Sonar–Optical Multimodal Large… 8.0 前25% 方法研究 #音频理解 14. CoSTALA: Compositional Spatio-Temporal Audio-Language… 7.9 前25% 方法研究 #空间音频 15. Don’t Just Listen, Try Planning: Graph-based Retrieval… 7.8 前25% 方法研究 #音频理解 16. On the Robustness of Audio Deepfake Detection under… 7.7 前25% 方法研究 #音频伪造检测 17. Speech-to-SOAP: End-to-End Summarization of Medical… 7.7 前25% 系统技术报告 #音频理解 18. Array-Agnostic Ambisonics Encoding via Diffusion… 7.5 前25% 方法研究 #空间音频 19. OmniJudge or OmniBias? Diagnosing Multimodal Judges… 7.4 前50% 数据集与基准 #音频质量评估 20. Task-disentangled Low-Rank Adaptation for Versatile… 7.4 前50% 方法研究 #音视频理解 21. Anatomy of a Scam Call: What 10,000 real scam and spam… 7.3 前50% 应用研究 #语音交互 22. Preference Optimization for Non-Verbal Vocalization… 7.1 前50% 方法研究 #语音合成 23. One Timeline, Many Renderings: A Wolfram Language… 7.0 前50% 系统技术报告 #音乐生成 24. Visually-Guided Spatial Audio Generation for… 6.9 前50% 方法研究 #空间音频 25. Benchmarking LLM Judges for Voice-Agent Evaluation… 6.6 前50% 数据集与基准 #语音交互 26. Investigating voiced and unvoiced regions of speech… 6.4 前50% 方法研究 #语音伪造检测 📋 论文列表 🥇 EmoTra-TTS: Smooth Intra-Utterance Emotion Transitions for Speech Synthesis 9.6/10 | 创新 1.7/2 | 严谨 1.4/1.5 | 实验 1.5/1.5 | 清晰 0.9/1 | 影响 1.2/1.5 | 开源 1.2/1.5 | 复现 0.4/0.5 | 工程 1.3/1.5 ...
语音/音乐/音频论文速递 2026-08-25
语音/音乐/音频论文速递 2026-08-25 共分析 46 篇论文 ⚡ 今日概览 ✅ 筛选入选 46 篇 → 🔬 深度分析完成 🏷️ 热门方向 方向 数量 分布 #语音识别 7 篇 ███████ #音视频理解 4 篇 ████ #语音交互 3 篇 ███ #语音增强 3 篇 ███ #语音情感识别 3 篇 ███ #音频事件检测 3 篇 ███ #音频理解 3 篇 ███ #语音合成 2 篇 ██ 📊 论文评分排行榜(46 篇,按分数降序) 排名 论文 总分 分档 文档类型 主任务 🥇 TLive-Omni: An Omni-Modal Understanding Model for E… 10.0 前10% 系统技术报告 #音视频理解 🥈 sanoTTS: The Smallest Real-Time Neural TTS on a… 10.0 前10% 系统技术报告 #语音合成 🥉 Do Spoken Language Models Hear Speech as They Read… 10.0 前10% 方法研究 #语音交互 4. LipsAM: Lipschitz-continuous Neural Networks for… 10.0 前10% 理论研究 #音频修复 5. PolyChirp: Multi-Species Birdsong Classification Using… 10.0 前10% 系统技术报告 #音频分类 6. EchoWM: Open and Enterable Omnimodal World Models 10.0 前10% 系统技术报告 #音视频生成 7. Simulation-to-Real First-Break Segmentation for… 9.8 前10% 应用研究 #音频事件检测 8. A Computationally Efficient Likelihood Approximation… 9.8 前10% 方法研究 #声源定位 9. A Factorial Ablation of a Speech-to-SFT Pipeline… 9.7 前10% 应用研究 #语音交互 10. Self-Supervised Speech Representations Track Spoken… 9.7 前10% 应用研究 #语音属性识别 11. AudioWorldSim: Realistic Binaural Audio Datasets For… 9.7 前10% 系统技术报告 #空间音频 12. AT-ADD: A Benchmark and Challenge for Robust and All… 9.6 前10% 数据集与基准 #音频伪造检测 13. WnW: Waxing-and-Waning KV Cache for Long-Form Speech… 9.4 前10% 方法研究 #语音交互 14. Better Retrieval, Worse Robustness: How Multi-hop RAG… 9.1 前10% 应用研究 #音频理解 15. Training DeepFilterNet with Accurate Room Acoustic… 9.0 前10% 应用研究 #语音增强 16. TurboBias 2.0: Streaming Context-Biasing for… 9.0 前10% 系统技术报告 #语音识别 17. FlowSep 2: Self-Supervised Flow Matching for Language… 9.0 前10% 方法研究 #音频分离 18. Pre-Decoding Acoustic Triage for Budgeted Vision… 9.0 前10% 方法研究 #音视频理解 19. Do SpeechLMs Hear Their Own Opinions? Diagnosing and… 8.9 前25% 方法研究 #语音情感识别 20. MRMAD: A Multi-Round Multi-Audio Benchmark for… 8.9 前25% 数据集与基准 #音频质量评估 21. Unsupervised Speech Recognition at the Syllable Level 8.9 前25% 方法研究 #语音识别 22. Towards Actionable Surgical Team Dynamics: from… 8.9 前25% 数据集与基准 #音视频理解 23. Long-Horizon Audio-Visual Generation for Persistent… 8.9 前25% 系统技术报告 #音视频生成 24. Do Time-Series Foundation Models Pay Off for… 8.8 前25% 应用研究 #音频事件检测 25. MusPyExpress: Extending MusPy with Enhanced Expression… 8.7 前25% 系统技术报告 #音乐理解 26. Adaptive Hierarchical Representation Alliance for… 8.6 前25% 方法研究 #语音情感识别 27. Vibrato Matching for Modulation Control and Blending… 8.5 前25% 方法研究 #音频生成 28. Spiking Neural Networks for Energy-Efficient Object… 8.5 前25% 应用研究 #音频事件检测 29. Development and Feasibility Evaluation of an Edge AI… 8.4 前25% 应用研究 #语音识别 30. Dual-Scale State-Space Modeling with Speaker-Wise… 8.4 前25% 方法研究 #语音情感识别 31. μNet: Ultra-Low-Memory and Low-Complexity Speech… 8.3 前25% 方法研究 #语音增强 32. Cross-Subject Generalization in Decoding Perceived… 8.3 前25% 方法研究 #语音识别 33. Reasoning-Oriented Post-Training and Inference-Time… 8.3 前25% 应用研究 #音频理解 34. Multi-Modal Semantic Expansion with Constrained LLM… 8.3 前25% 系统技术报告 #音乐推荐 35. DAMOS: Learning Distortion-Aware Speech Quality… 8.2 前25% 方法研究 #语音质量评估 36. Multi-Task Learning for Non-Canonical Phoneme… 8.2 前25% 方法研究 #语音识别 37. MetaSICL: Globalizing Auditory LLMs for Underserved… 8.0 前25% 方法研究 #音频理解 38. AudioNoisePrints: Model-free audio watermarking using… 7.9 前25% 方法研究 #音频水印 39. Building and Evaluating a Synthetic Bengali Speech… 7.6 前25% 数据集与基准 #语音合成 40. SlimDiffuSE: Towards Efficient Diffusion-Based Speech… 7.6 前25% 方法研究 #语音增强 41. DiaScriber: A Speech LLM for Joint Diarization and… 7.3 前50% 方法研究 #语音识别 42. A Regularized Block Diagonal RLS Algorithm for… 7.2 前50% 方法研究 #回声消除 43. Separating Voice from Age in COPD Screening 7.2 前50% 应用研究 #语音属性识别 44. Mitigating Speaker Leakage in Cascaded Multi-talker… 7.1 前50% 方法研究 #语音识别 45. Motion-Aware Reasoning from Speech to Mask Tracks… 7.1 前50% 系统技术报告 #音视频理解 46. Humanoid Musical Robots as Experimental Interfaces for… 5.9 前50% 应用研究 #音视频交互 📋 论文列表 🥇 TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming 10.0/10 | 创新 1.7/2 | 严谨 1.4/1.5 | 实验 1.5/1.5 | 清晰 1/1 | 影响 1.5/1.5 | 开源 1.5/1.5 | 复现 0.4/0.5 | 工程 1/1.5 ...
语音/音乐/音频论文速递 2026-08-21
语音/音乐/音频论文速递 2026-08-21 共分析 18 篇论文 ⚡ 今日概览 ✅ 筛选入选 18 篇 → 🔬 深度分析完成 🏷️ 热门方向 方向 数量 分布 #音频理解 3篇 ███ #语音交互 2篇 ██ #语音伪造检测 2篇 ██ #语音识别 2篇 ██ #音乐生成 2篇 ██ #助听器 1篇 █ #语音情感识别 1篇 █ #音乐检索 1篇 █ 📊 论文评分排行榜(18 篇,按分数降序) 排名 论文 总分 分档 文档类型 主任务 🥇 Listening Forward: Next Patch Embedding Prediction… 8.8分 前25% 方法研究 #音频理解 🥈 Towards Quantifying Benchmark Optimization in ASR… 8.5分 前25% 方法研究 #语音识别 🥉 \(TCP_α\): Margin-Controlled Confidence estimation for… 8.4分 前25% 方法研究 #音乐理解 4. VA-Judger: Reward Modeling from Human Preference… 8.3分 前25% 方法研究 #音视频生成 5. Unified Music Identification for Tracks and Versions 8.2分 前25% 数据集与基准 #音乐检索 6. Fourier is Frontier: Frequency-Aware Autoencoding for… 7.5分 前25% 方法研究 #音乐生成 7. Explainability by Design: Structured Kolmogorov-Arnold… 7.5分 前25% 方法研究 #语音伪造检测 8. Represented but Ignored: A Causal Account of Prosodic… 7.2分 前50% 方法研究 #音频理解 9. A Speech Corpus for Mizo Automatic Speech Recognition… 6.7分 前50% 数据集与基准 #语音识别 10. Hear2Act: Benchmarking When Prosody Should Change What… 6.4分 前50% 数据集与基准 #语音交互 11. DAVSS: Distilled Audio-Visual State Space Models 6.4分 前50% 方法研究 #音视频理解 12. A Resource-Efficient CNN-Based EEG Auditory Attention… 6.2分 前50% 系统技术报告 #助听器 13. Tracking the Trend in How Speech Synthesizers Deceive… 6.0分 前50% 应用研究 #语音伪造检测 14. Does Mapping Non-Maximal Probabilities to GMM… 5.6分 前50% 方法研究 #音频理解 15. MultiVerse: A Creator-Centered Approach to Steering… 5.4分 后50% 应用研究 #音乐生成 16. Robust Incomplete Multimodal Sentiment Analysis via… 5.4分 后50% 方法研究 #语音情感识别 17. Does Listening Matter? Backchanneling and Nodding in… 5.2分 后50% 应用研究 #语音交互 18. Dancing Through Soundscapes: Designing a Low-Cost… 5.2分 后50% 应用研究 #音频交互 📋 论文列表 🥇 Listening Forward: Next Patch Embedding Prediction Enables Scalable Audio Learners 8.8/10 | 创新 1.5/2 | 严谨 1.3/1.5 | 实验 1.4/1.5 | 清晰 0.8/1 | 影响 1.2/1.5 | 开源 1.2/1.5 | 复现 0.4/0.5 | 工程 1/1.5 ...