论文解读

DrVoice: Parallel Speech-Text Voice Conversation Model via Dual-Resolution Speech Representations

语音对话系统 | 9.5/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4530 字 阅读 →
论文解读

EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models

基准测试 | 7.0/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4360 字 阅读 →
论文解读

End-to-end Listen, Look, Speak and Act

语音对话系统 | 8.5/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5323 字 阅读 →
论文解读

From Text to Talk: Audio-Language Model Needs Non-Autoregressive Joint Training

语音对话系统 | 8.5/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5679 字 阅读 →
论文解读

Human or Machine? A Preliminary Turing Test for Speech-to-Speech Interaction

语音对话系统 | 7.5/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4701 字 阅读 →
会议任务专题

ICLR 2026 - 语音对话系统

共 8 篇 ICLR 2026 语音对话系统 方向论文

 · 更新于 2026-09-24 · 约 23 分钟 · 11507 字 阅读 →
论文解读

NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching

多模态模型 | 8.0/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5387 字 阅读 →
论文解读

ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech Interaction

语音对话系统 | 8.0/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4678 字 阅读 →
论文解读

RoboOmni: Proactive Robot Manipulation in Omni-modal Context

机器人操作 | 7.5/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5562 字 阅读 →
论文解读

STITCH: Simultaneous Thinking and Talking with Chunked Reasoning for Spoken Language Models

语音对话系统 | 8.5/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4500 字 阅读 →
论文解读

Towards True Speech-to-Speech Models Without Text Guidance

语音对话系统 | 7.5/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5080 字 阅读 →
论文解读

Can Speech LLMs Think while Listening?

语音对话系统 | 7.5/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4627 字 阅读 →
论文解读

Closing the Gap Between Text and Speech Understanding in LLMs

语音对话系统 | 7.5/10

 · 更新于 2026-09-24 · 约 17 分钟 · 8119 字 阅读 →
论文解读

DrVoice: Parallel Speech-Text Voice Conversation Model via Dual-Resolution Speech Representations

语音对话系统 | 8.5/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5139 字 阅读 →
论文解读

EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models

基准测试 | 8.5/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5102 字 阅读 →
论文解读

End-to-end Listen, Look, Speak and Act

语音对话系统 | 8.5/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6653 字 阅读 →
论文解读

From Text to Talk: Audio-Language Model Needs Non-Autoregressive Joint Training

语音对话系统 | 7.5/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5535 字 阅读 →
论文解读

Human or Machine? A Preliminary Turing Test for Speech-to-Speech Interaction

语音对话系统 | 9.0/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4643 字 阅读 →
论文解读

MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction

语音对话系统 | 8.5/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5987 字 阅读 →
论文解读

NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching

多模态模型 | 8.5/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5125 字 阅读 →