论文解读

TTA: Transcribe, Translate and Alignment for Cross-Lingual Speech Representation

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5261 字 阅读 →
论文解读

Universr: Unified and Versatile Audio Super-Resolution Via Vocoder-Free Flow Matching

音频超分辨率 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4168 字 阅读 →
论文解读

Unseen but Not Unknown: Using Dataset Concealment to Robustly Evaluate Speech Quality Estimation Models

语音质量评估 | 8.3/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4293 字 阅读 →
论文解读

Utilizing Information Theoretic Approach to Study Cochlear Neural Degeneration

生物声学 | 6.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3850 字 阅读 →
论文解读

V2A-DPO: Omni-Preference Optimization for Video-To-Audio Generation

视频到音频生成 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4613 字 阅读 →
论文解读

Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models

音频问答 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3527 字 阅读 →
论文解读

WAV2LEV: Predicting Levenshtein Edit Operation Sequences For Fine-Grained Estimation of Automatic Speech Recognition Error

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4224 字 阅读 →
论文解读

When Noise Lowers the Loss: Rethinking Likelihood-Based Evaluation in Music Large Language Models

音乐生成 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4847 字 阅读 →
论文解读

When Silence Matters: The Impact of Irrelevant Audio on Text Reasoning in Large Audio-Language Models

模型评估 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4322 字 阅读 →
论文解读

When Voice Matters: A Controlled Study of Audio LLM Behavior in Clinical Decision-Making

模型评估 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3980 字 阅读 →
论文解读

Why Do Speech Language Models Fail to Generate Semantically Coherent Outputs? A Modality Evolving Perspective

语音生成 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4296 字 阅读 →
论文解读

Z-Scores: A Metric for Linguistically Assessing Disfluency Removal

模型评估 | 6.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4710 字 阅读 →
论文解读

All That Glitters Is Not Audio: Rethinking Text Priors and Audio Reliance in Audio-Language Evaluation

音频问答 | 6.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4036 字 阅读 →
论文解读

Comparison of sEMG Encoding Accuracy Across Speech Modes Using Articulatory and Phoneme Features

语音生物标志物 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4365 字 阅读 →
论文解读

Explainable AI in Speaker Recognition -- Making Latent Representations Understandable

说话人识别 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3672 字 阅读 →
论文解读

Audio Video Verbal Analysis (AVVA) for Capturing Classroom Dialogues

音频问答 | 6.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4018 字 阅读 →
论文解读

Identifying and typifying demographic unfairness in phoneme-level embeddings of self-supervised speech recognition models

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4567 字 阅读 →
论文解读

Transformer-Based Rhythm Quantization of Performance MIDI Using Beat Annotations

音乐信息检索 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4294 字 阅读 →
论文解读

"This Wasn't Made for Me": Recentering User Experience and Emotional Impact in the Evaluation of ASR Bias

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 4 分钟 · 1997 字 阅读 →
论文解读

AUDITA: A New Dataset to Audit Humans vs. AI Skill at Audio QA

音频问答 | 6.5/10

 · 更新于 2026-09-25 · 约 5 分钟 · 2444 字 阅读 →