论文解读

Spatio-Temporal Audio Language Modeling for Dynamic Sound Sources

音频问答 | 6.9/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5968 字 阅读 →
论文解读

The Holistic Storage of Verb+Up Phrases in Text-based and Audio-based Language Models

语音识别 | 8.2/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5775 字 阅读 →
论文解读

The Perceived Fragility of Explanations in Audio Models: Manipulation of Attribution with Unchanged Predictions

The Perceived Fragility of Explanations in Audio Models: Manipulation of Attribution with Unchanged Predictions

 · 更新于 2026-09-07 · 约 11 分钟 · 5388 字 阅读 →
论文解读

Unsupervised Approaches for Global Prosodic Embedding Extraction

语音合成 | 7.8/10

 · 更新于 2026-09-07 · 约 13 分钟 · 6086 字 阅读 →
论文解读

Who Spoke When in Multi-Conversation: Target Speaker Tagging Task and Benchmark

说话人识别 | 8.6/10

 · 更新于 2026-09-07 · 约 10 分钟 · 4749 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-15

共分析 26 篇语音/AI 论文

 · 更新于 2026-09-07 · 约 72 分钟 · 35999 字 阅读 →
论文解读

A Dual-Mode Faust-to-CLAP Compilation System

A Dual-Mode Faust-to-CLAP Compilation System

 · 更新于 2026-09-07 · 约 10 分钟 · 5003 字 阅读 →
论文解读

Adaptive Turn-Taking for Real-time Multi-Party Voice Agents

数据增强 | 6.7/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5916 字 阅读 →
论文解读

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation

音频生成 | 9/10

 · 更新于 2026-09-07 · 约 14 分钟 · 6874 字 阅读 →
论文解读

Balancing ASR and diarization in end-to-end LLMs for multi-talker speech recognition

语音识别 | 7.1/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5831 字 阅读 →
论文解读

BASENet: Band-Adapted Speech Enhancement Network with Cross-Band Attention

语音增强 | 7.5/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5907 字 阅读 →
论文解读

Decoding Insect Song: A Multitask Semisupervised Orthoptera Bioacoustic Classifier

音频分类 | 8.7/10

 · 更新于 2026-09-07 · 约 14 分钟 · 6708 字 阅读 →
论文解读

Dolph2Vec: Self-Supervised Representations of Dolphin Vocalizations

音频分类 | 7.2/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5440 字 阅读 →
论文解读

Emo-LiPO: Listwise Preference Optimization for Fine-Grained Emotion Intensity Control in LLM-based Text-to-Speech

语音合成 | 9.3/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5378 字 阅读 →
论文解读

Endpoint Anticipation for Low-Latency Spoken Dialogue

多任务学习 | 8.2/10

 · 更新于 2026-09-07 · 约 9 分钟 · 4326 字 阅读 →
论文解读

From Tokens to Faces: Investigating Discrete Speech Representations for 3D Facial Animation

语音合成 | 7.9/10

 · 更新于 2026-09-07 · 约 13 分钟 · 6037 字 阅读 →
论文解读

Generating Training Targets for Real-World Speech Enhancement via Close-to-Distant Microphone Projection

语音增强 | 6.4/10

 · 更新于 2026-09-07 · 约 10 分钟 · 4533 字 阅读 →
论文解读

Generative Modeling of Bach-Style Symbolic Music: A Comparative Study of Autoregressive, Latent-Variable, and Adversarial Approaches

音乐生成 | 5.7/10

 · 更新于 2026-09-07 · 约 10 分钟 · 4808 字 阅读 →
论文解读

Leveraging Audio-LLMs to Filter Speech-to-Speech Training Data

语音翻译 | 8.4/10

 · 更新于 2026-09-07 · 约 13 分钟 · 6061 字 阅读 →
论文解读

Low-Latency Real-Time Audio Game Commentary System via LLM-Based Parallel Text Generation

语音合成 | 7.9/10

 · 更新于 2026-09-07 · 约 10 分钟 · 4803 字 阅读 →