论文解读

When Synthetic Speech Is All You Have: Better Call GRPO

语音识别 | 7.7/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8468 字 阅读 →
论文解读

When Synthetic Speech Is All You Have: Better Call GRPO

语音识别 | 7.8/10

 · 更新于 2026-09-06 · 约 18 分钟 · 8582 字 阅读 →
论文解读

Why Do You Say It Like That? A Phoneme-Level Framework for Explainable Speech Deepfake Detection

语音伪造检测 | 6.5/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6461 字 阅读 →
论文解读

Why Do You Say It Like That? A Phoneme-Level Framework for Explainable Speech Deepfake Detection

语音伪造检测 | 6.0/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7622 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-07-10

共分析 19 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 64 分钟 · 31889 字 阅读 →
论文解读

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts

语音情感识别 | 6.9/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8349 字 阅读 →
论文解读

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs

语音识别 | 7/10

 · 更新于 2026-09-06 · 约 20 分钟 · 9935 字 阅读 →
论文解读

Decoupling Conversational Dynamics in Full-Duplex Spoken Models through Reinforcement Learning

语音交互 | 8.2/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7653 字 阅读 →
论文解读

EscFOA: Enhancing Spatial Learning for Visually Impaired Learners via Generative Spatial Audio in 360-Degree Educational Environments

教育 | 2.8/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6587 字 阅读 →
论文解读

Extending Xenakis: From Architectural Geometry to Sonification of the Philips Pavilion

音乐生成 | 5.6/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6884 字 阅读 →
论文解读

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs

语音识别 | 7.3/10

 · 更新于 2026-09-06 · 约 18 分钟 · 8554 字 阅读 →
论文解读

MADB: A Large-Scale Music Aesthetics Dataset with Professional and Multi-Dimensional Annotations

音乐理解 | 8.1/10

 · 更新于 2026-09-06 · 约 25 分钟 · 12029 字 阅读 →
论文解读

MMGenre: Benchmarking Singing Voice Synthesis across Multiple Musical Genres

基准测试 | 8.3/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6401 字 阅读 →
论文解读

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders

语音活动检测 | 6.7/10

 · 更新于 2026-09-06 · 约 20 分钟 · 9878 字 阅读 →
论文解读

Rag Classification of Tagore Songs using Symbolic Music Notation and Novel Weighted Distance Measures

音乐理解 | 3/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4733 字 阅读 →
论文解读

Text-Independent Speaker Verification Using Discrete Audio Tokens

说话人验证 | 5.2/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7091 字 阅读 →
论文解读

Transformer-based segmentation of prosodic boundaries in Brazilian Portuguese

语音识别 | 4/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7485 字 阅读 →
论文解读

UBG-Net: An Uncertainty-aware Bayesian Gating Network for Robust Audio-Visual Speech Recognition

语音识别 | 7.1/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6172 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-07-09

共分析 13 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 47 分钟 · 23491 字 阅读 →
论文解读

BlueMagpie-TTS: A Token-Efficient Tokenizer, Language Model, and TTS for Taiwanese-Accent Code-Switching Speech

语音合成 | 6.5/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6017 字 阅读 →