论文解读

Separating Decision-Rule Misalignment from Readout-Coverage Limitations in Speech Language Models

语音情感识别 | 6.8/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7558 字 阅读 →
论文解读

A Study of ASR Adaptation and Representation Dimensionality Reduction in Persian Speech Emotion Recognition Using Whisper

语音情感识别 | 5.7/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4924 字 阅读 →
论文解读

How to Recognize New Words: A Comparison Between Context Biasing Methods and Speech LLMs

语音识别 | 6.4/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7408 字 阅读 →
论文解读

Breaking the Curse of Multilinguality in Many-to-Many Speech-to-Text Translation via a Resource-Aware Mixture of Speech Encoders

语音翻译 | 8.2/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7762 字 阅读 →
论文解读

HyPASE: Hyperbolic Geometry for Parameter-Efficient Speech Emotion Fine-Tuning Framework for Large Audio-Language Models

语音情感识别 | 7.2/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7528 字 阅读 →
论文解读

Language-Specialized Multi-Teacher On-Policy Distillation for Multilingual LLM-Based ASR

语音识别 | 6.9/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6440 字 阅读 →
论文解读

JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents

语音合成 | 6.5/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5563 字 阅读 →
论文解读

Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm

语音合成 | 7.0/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7986 字 阅读 →
论文解读

Kutti AI: A Voice-First, Offline-Capable Learning Companion with Real-Time Struggle Detection for Visually-Impaired Children

语音交互 | 5.5/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4866 字 阅读 →
论文解读

Safeguards for Speech2Speech LLM-Assistants: A Case Study in Automotive Applications

语音交互 | 6.5/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4430 字 阅读 →
论文解读

Toward Generalizable Cognitive Impairment Detection with Speech-Based Multimodal Large Language Models

语音情感识别 | 7.0/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6446 字 阅读 →
论文解读

Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models

语音交互 | 7.6/10

 · 更新于 2026-09-24 · 约 18 分钟 · 8635 字 阅读 →
论文解读

Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing

语音识别 | 7.1/10

 · 更新于 2026-09-24 · 约 18 分钟 · 8698 字 阅读 →
论文解读

Robust Summarization of Doctor-Patient Conversations: TalTech Systems for the Beyond Transcription Challenge

语音交互 | 6.3/10

 · 更新于 2026-09-24 · 约 18 分钟 · 8701 字 阅读 →
论文解读

Do LLMs Need Architectural Changes for Simultaneous Speech Translation? A Prefix-to-Prefix Data Driven Approach

语音翻译 | 5.7/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7514 字 阅读 →
论文解读

Breaking the Quality--Intelligibility Trade-off in Streaming Target Speaker Extraction via Deep-Feature-Anchored Preference Optimization

语音分离 | 6.3/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7612 字 阅读 →
论文解读

GigaAM Multilingual: Foundation Model for Underrepresented Languages

语音识别 | 8.1/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6967 字 阅读 →
论文解读

COALA: Robust Contextualized Speech-augmented Language Modeling for ASR via Contrastive Regularizer and Biasing Score Estimation

语音识别 | 7.5/10

 · 更新于 2026-09-24 · 约 19 分钟 · 9228 字 阅读 →
论文解读

Diarization-Guided Qwen-ASR Adaptation for Multilingual Two-Speaker Conversational Speech

语音识别 | 5.7/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6538 字 阅读 →
论文解读

Diarization-Guided Qwen-ASR Adaptation for Multilingual Two-Speaker Conversational Speech

语音识别 | 7.0/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7440 字 阅读 →