论文解读

Numerical and perceptual validity of synthetic Head-Related Transfer Functions at scale

声源定位 | 7.6/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8223 字 阅读 →
论文解读

The Null Token Knows: Reducing Message-Free Hallucination in ASR and NMT

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7035 字 阅读 →
论文解读

Using the Mimi codec for metalinguistic representations

语音编码 | 4.7/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5440 字 阅读 →
论文解读

What Makes a Good Layer? Assessing the Layer-Wise Intrinsic Properties of Music Foundation Models

音乐理解 | 7.7/10

 · 更新于 2026-09-25 · 约 16 分钟 · 8007 字 阅读 →
论文解读

Measuring Fairness in Large Audio Language Models via Semantic-Aware Bias Estimation

语音识别 | 5.6/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7337 字 阅读 →
论文解读

Sensor-Driven Mission Synthesis for UAV/UGV Swarms: A TB-CSPN Coordination Architecture with Hardware-Enforced Safety

音视频理解 | 4.4/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6424 字 阅读 →
论文解读

Trajectory Dynamics in Self-Supervised Learning Latent Space for Audio Deepfake Detection

语音伪造检测 | 7.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5887 字 阅读 →
论文解读

Why Performance Metrics Overpromise in Auditory Attention Decoding: an Information-Theoretic Reappraisal

语音交互 | 6.9/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6654 字 阅读 →
论文解读

Evaluating Pre-trained Speech Encoders for Spontaneous Speech Detection and Out of Domain Synthetic Speech Generalisation in Indic Languages

语音伪造检测 | 6.0/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6220 字 阅读 →
论文解读

HybridSB-MoE: Dual-Domain Schrödinger Bridges with Scene-Adaptive Expert Routing for Speech Enhancement

语音增强 | 5.9/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7143 字 阅读 →
论文解读

Deep Learning Based Relative Transfer Matrix Estimation for Multiple Sources and Multiple Microphones

语音增强 | 5.1/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6920 字 阅读 →
论文解读

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping

音乐生成 | 6.8/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8231 字 阅读 →
论文解读

MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching

音频生成 | 8.0/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7722 字 阅读 →
论文解读

ASR-Roundtrip Evaluation Can Mask Context- and Convention-Dependent Reading Errors in Chinese News TTS

语音合成 | 6.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6433 字 阅读 →
论文解读

DuplexWorld: Can voice agents help you get through the day?

语音交互 | 8.5/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8820 字 阅读 →
论文解读

In Defense of Using Worst-case Privacy Disclosure as Privacy Evaluation Metric of Voice Anonymization

说话人验证 | 7.1/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5631 字 阅读 →
论文解读

Measuring Cross-Cultural Style Diffusion Through Era Classification: US and Korean Popular Music

音乐理解 | 8.3/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7089 字 阅读 →
论文解读

TimeRoute: Time-Aware Modality Routing and Diffusion for Multi-Modal Recommendation

扩散模型 | 6.5/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9515 字 阅读 →
论文解读

Training Set Synthesis for Bioacoustic Denoising: A Case Study With Mice

语音增强 | 8.0/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8643 字 阅读 →
论文解读

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction

语音交互 | 6.4/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5797 字 阅读 →