论文解读

听错一个字就全盘皆输:SpeechGym 把语音智能体的失败变成可训练的梯度

语音交互 | 7.6/10

 · 更新于 2026-09-24 · 约 17 分钟 · 8505 字 阅读 →
论文解读

克隆与匿名本是一体两面:把 27k 小时的 TTS 直接当成隐私滤镜

语音转换 | 8.0/10

 · 更新于 2026-09-24 · 约 17 分钟 · 8455 字 阅读 →
论文解读

Relative Time Intervals Representation for Word-level Timestamping with Masked Training

语音识别 | 8.6/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6295 字 阅读 →
论文解读

Speech-to-SOAP: End-to-End Summarization of Medical Dialogues: KIT@BeTraC 2026

音频理解 | 7.7/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6505 字 阅读 →
论文解读

DiaScriber: A Speech LLM for Joint Diarization and Transcription in Multi-Speaker Scenarios

语音识别 | 7.3/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4845 字 阅读 →
论文解读

Do SpeechLMs Hear Their Own Opinions? Diagnosing and Mitigating Previous-Belief Contamination in Streaming Emotion Understanding

语音情感识别 | 8.9/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4535 字 阅读 →
论文解读

Do Spoken Language Models Hear Speech as They Read Text? Bridging Structural Gaps Between Speech and Text

语音交互 | 10.0/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4838 字 阅读 →
论文解读

WnW: Waxing-and-Waning KV Cache for Long-Form Speech LLMs

语音交互 | 9.4/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5062 字 阅读 →
论文解读

Aslema at NADI 2026: Augmentation through Fewshot for SLU

语音交互 | 8.0/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4468 字 阅读 →
论文解读

A survey of AI-generated voices and their detection

语音伪造检测 | 6.1/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4920 字 阅读 →
论文解读

DuplexGen: Decoupling Content, Timing, and Acoustics for Synthetic Dialogue Speech

语音合成 | 6.1/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6343 字 阅读 →
论文解读

Measuring Fairness in Large Audio Language Models via Semantic-Aware Bias Estimation

语音识别 | 5.6/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7337 字 阅读 →
论文解读

VoiceChat-TTS: A Low-Latency Continuous Speech Synthesis Model for Interactive Agents

语音合成 | 8.2/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7721 字 阅读 →
论文解读

Rethinking Language Model-Based Generative Speech Enhancement in the Latent Space of a Neural Audio Codec

语音增强 | 7.1/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5494 字 阅读 →
论文解读

The SLT 2026 SmartGlasses Challenge: Benchmarking Egocentric Multi-Talker Speech Recognition and Understanding with Audio-Language Models

语音识别 | 7.5/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6748 字 阅读 →
论文解读

Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions

语音质量评估 | 7.1/10

 · 更新于 2026-09-24 · 约 17 分钟 · 8060 字 阅读 →
论文解读

Whisper-Aware LLM: Self-Supervised Uncertainty Learning for Robust Whispered Speech Recognition

语音识别 | 6.1/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6535 字 阅读 →
论文解读

Speaker Role and Language Diarization for Analyzing Multilingual Interviews for Language Proficiency of Older Adults

说话人日志 | 6.3/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6762 字 阅读 →
论文解读

SraVaani 1.0: Scaling Inclusive Speech Recognition for Indic Languages

语音识别 | 7.9/10

 · 更新于 2026-09-24 · 约 17 分钟 · 8512 字 阅读 →
论文解读

VoxZip: Semantic-Anchored Temporal KV Cache Compression for Long-Context Audio Inference

音频理解 | 7.7/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6697 字 阅读 →