论文解读

Mispronunciation Detection and Diagnosis Without Model Training: A Retrieval-Based Approach

语音评估 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4234 字 阅读 →
论文解读

Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4158 字 阅读 →
论文解读

MNV-17: A High-Quality Performative Mandarin Dataset for Nonverbal Vocalization Recognition in Speech

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3597 字 阅读 →
论文解读

MSF-SER: Enriching Acoustic Modeling with Multi-Granularity Semantics for Speech Emotion Recognition

语音情感识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5146 字 阅读 →
论文解读

MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-Token Prediction

语音翻译 | 8.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5619 字 阅读 →
论文解读

Optimizing Speech Language Models for Acoustic Consistency

语音合成 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4747 字 阅读 →
论文解读

PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models

语音对话系统 | 8.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4703 字 阅读 →
论文解读

Plug-and-Play Emotion Graphs for Compositional Prompting in Zero-Shot Speech Emotion Recognition

语音情感识别 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4538 字 阅读 →
论文解读

Principled Coarse-Grained Acceptance For Speculative Decoding In Speech

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3911 字 阅读 →
论文解读

Proficiency-Aware Adaptation and Data Augmentation for Robust L2 ASR

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3992 字 阅读 →
论文解读

Reducing Prompt Sensitivity in LLM-Based Speech Recognition Through Learnable Projection

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4818 字 阅读 →
论文解读

Reference-Aware SFM Layers for Intrusive Intelligibility Prediction

语音评估 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4990 字 阅读 →
论文解读

Relative Time Intervals Representation For Word-Level Timestamping With Masked Training

语音识别 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4823 字 阅读 →
论文解读

Revisiting Direct Speech-to-Text Translation with Speech LLMS: Better Scaling than Cot Prompting?

语音翻译 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4766 字 阅读 →
论文解读

RLBR: Reinforcement Learning with Biasing Rewards for Contextual Speech Large Language Models

语音识别 | 8.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3657 字 阅读 →
论文解读

Scaling Spoken Language Models with Syllabic Speech Tokenization

语音理解 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4144 字 阅读 →
论文解读

SED: Structural Entropy Based Speech Discretization for Discrete Token-Based ASR

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4993 字 阅读 →
论文解读

Session-Level Spoken Language Assessment with A Multimodal Foundation Model Via Multi-Target Learning

语音评估 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4959 字 阅读 →
论文解读

SLM-SS: Speech Language Model for Generative Speech Separation

语音分离 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5237 字 阅读 →
论文解读

SLM-TTA: A Framework for Test-Time Adaptation of Generative Spoken Language Models

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4556 字 阅读 →