论文解读

SmartDJ: Declarative Audio Editing with Audio Language Model

音频生成 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4459 字 阅读 →
论文解读

Speech World Model: Causal State–Action Planning with Explicit Reasoning for Speech

语音情感识别 语音对话系统 | 9.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5094 字 阅读 →
论文解读

Speech-to-LaTeX: New Models and Datasets for Converting Spoken Equations and Sentences

语音识别 | 8.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4797 字 阅读 →
论文解读

SpeechJudge: Towards Human-Level Judgment for Speech Naturalness

模型评估 | 8.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5451 字 阅读 →
论文解读

TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling

语音对话系统 | 8.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3884 字 阅读 →
论文解读

Towards True Speech-to-Speech Models Without Text Guidance

语音对话系统 | 9.1/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4831 字 阅读 →
论文解读

UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice

语音翻译 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5273 字 阅读 →
论文解读

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Video

多模态模型 | 7.0/10

 · 更新于 2026-09-25 · 约 7 分钟 · 3342 字 阅读 →
论文解读

VowelPrompt: Hearing Speech Emotions from Text via Vowel-level Prosodic Augmentation

语音情感识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5344 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-02

共分析 4 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 15 分钟 · 7026 字 阅读 →
论文解读

BUT System Description for CHiME-9 MCoRec Challenge

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5308 字 阅读 →
论文解读

DM-ASR: Diarization-aware Multi-speaker ASR with Large Language Models

说话人识别 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4581 字 阅读 →
论文解读

Do Sparse Autoencoders Capture Concept Manifolds?

可解释性 | 7.0/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7300 字 阅读 →
论文解读

Few-Shot Accent Synthesis for ASR with LLM-Guided Phoneme Editing

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5168 字 阅读 →
论文解读

MCPHunt: An Evaluation Framework for Cross-Boundary Data Propagation in Multi-Server MCP Agents

模型评估 | 7.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6271 字 阅读 →
论文解读

Praxy Voice: Voice-Prompt Recovery + BUPS for Commercial-Class Indic TTS from a Frozen Non-Indic Base at Zero Commercial-Training-Data Cost

语音合成 | 8.0/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5558 字 阅读 →
论文解读

StarDrinks: An English and Korean Test Set for SLU Evaluation in a Drink Ordering Scenario

数据集 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3921 字 阅读 →
论文解读

Tatemae: Detecting Alignment Faking via Tool Selection in LLMs

大语言模型 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4820 字 阅读 →
论文解读

A Dataset of Robot-Patient and Doctor-Patient Medical Dialogues for Spoken Language Processing Tasks

语音对话系统 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3628 字 阅读 →
论文解读

A LLM-Driven Acoustic Semantic Enriched Framework for Underwater Acoustic Target Recognition

音频分类 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4601 字 阅读 →