论文解读

Explainable AI through the Lens of Material Agency: Enabling Musical Interface Design with Neural Audio Models

音乐生成 | 6.7/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6159 字 阅读 →
论文解读

Expose Your Disguise: Recovering Source Speaker Identity From Voice Conversion

对比学习 | 7.5/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6629 字 阅读 →
论文解读

Improving Zero-Shot Phonetic Classification through Language-Agnostic Articulatory Features

语音属性识别 | 5.3/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7242 字 阅读 →
论文解读

Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages

说话人日志 | 7.1/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8448 字 阅读 →
论文解读

Infinite Canons: Maximally Self-Similar Melodic Lines and Canons with Infinite Solutions

音乐生成 | 6.4/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6656 字 阅读 →
论文解读

JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents

音视频生成 | 6.2/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6160 字 阅读 →
论文解读

Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis

语音合成 | 6.9/10

 · 更新于 2026-09-06 · 约 21 分钟 · 10244 字 阅读 →
论文解读

Leveraging Gradient Reversal Loss and Multitask Learning for Datasets-Aware Audio Deepfake Detection

语音伪造检测 | 7.3/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6744 字 阅读 →
论文解读

Looking for Affect in Spontaneous Finnish Speech through Linguistic Interpretability

语音情感识别 | 5.3/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5152 字 阅读 →
论文解读

Low-Latency Turn-Taking via Context-Aware Preface Generation in a Real-World Dialogue Robot

语音交互 | 6.1/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5166 字 阅读 →
论文解读

Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers

语音合成 | 6.8/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7089 字 阅读 →
论文解读

Mind the Microphone Gap: Benchmarking Array Upsampling Strategies for Latent Acoustic Mapping

声源定位 | 6.6/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6635 字 阅读 →
论文解读

Modeling Stylistic Co-evolution in Symbolic Music Heritage Collections

音乐理解 | 4.9/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8057 字 阅读 →
论文解读

MoLGE: Mixture of Language Group Experts for Efficient Scaling of Massively Multilingual Speech Recognition

语音识别 | 6.7/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4906 字 阅读 →
论文解读

Multimodal Domain Generalization for Depression Detection: An Attention-Based BiLSTM Network with Domain-Adversarial Training

音频分类 | 6.4/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6423 字 阅读 →
论文解读

Music-Source-Separation-Training (MSST): A Unified Framework for Training and Evaluating Music Demixing Models

音乐源分离 | 5.7/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7537 字 阅读 →
论文解读

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation

音视频生成 | 7.1/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7731 字 阅读 →
论文解读

PathRIR: Physics-Guided Acoustic Path Selection and Late-Tail Compensation for Fast Room Impulse Response Simulation

空间音频 | 7.2/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6462 字 阅读 →
论文解读

Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm

语音合成 | 7.0/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7986 字 阅读 →
论文解读

Resource-Aware Topology Management for ISAC-Enabled TDOA Localization in IoUT Networks

声源定位 | 4.1/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5538 字 阅读 →