论文解读

Cross-Subject Semantic Decoding with Shared-Space Alignment for Generalized Neural Representation Learning

语音交互 | 6.3/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6689 字 阅读 →
论文解读

Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models

语音交互 | 7.6/10

 · 更新于 2026-09-24 · 约 18 分钟 · 8635 字 阅读 →
论文解读

End-to-End Markov State Sequence Learning for Auditory Attention Decoding

语音交互 | 8.3/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7112 字 阅读 →
论文解读

Harness TTS: Towards Context-Aware Expressive Speech Synthesis with Harness Layer

语音合成 | 6.2/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6532 字 阅读 →
论文解读

Robust Summarization of Doctor-Patient Conversations: TalTech Systems for the Beyond Transcription Challenge

语音交互 | 6.3/10

 · 更新于 2026-09-24 · 约 18 分钟 · 8701 字 阅读 →
论文解读

AnovaX: A Local, Multi-Agent Voice Assistant with LLM Planning, Typed Executors, and Adaptive Recovery

语音交互 | 4.8/10

 · 更新于 2026-09-24 · 约 21 分钟 · 10217 字 阅读 →
论文解读

Dialogs: a studio-quality expressive conversational Russian speech corpus for dialog assistants

语音合成 | 7.8/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6565 字 阅读 →
论文解读

Real-time Generation of Listener Nodding via Prediction of Kinematic Parameters for Avatar Dialogue Systems

语音交互 | 6.9/10

 · 更新于 2026-09-24 · 约 20 分钟 · 9767 字 阅读 →
论文解读

Casting Everything to Online API Services? A Survey of Integrating Localized Speech Recognition Models in Robotic Systems

语音识别 | 5.4/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7257 字 阅读 →
论文解读

Efficiently Adapting Spoken Language Models for the Singaporean Context

语音交互 | 6.5/10

 · 更新于 2026-09-24 · 约 17 分钟 · 8212 字 阅读 →
论文解读

Immersive Social Interaction with VR and LLM-Assisted Humanoids

语音交互 | 4.7/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5532 字 阅读 →
论文解读

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents

语音质量评估 | 8.7/10

 · 更新于 2026-09-24 · 约 19 分钟 · 9401 字 阅读 →
论文解读

Decoupling Conversational Dynamics in Full-Duplex Spoken Models through Reinforcement Learning

语音交互 | 8.2/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7653 字 阅读 →
论文解读

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs

语音交互 | 9.2/10

 · 更新于 2026-09-24 · 约 17 分钟 · 8435 字 阅读 →
论文解读

DuplexChat: Constructing Speaker-Separated Full-Duplex Dialogue Speech at Scale for Spoken Dialogue Language Modeling

语音交互 | 5.9/10

 · 更新于 2026-09-24 · 约 18 分钟 · 8798 字 阅读 →
论文解读

Metronome: Bound the Cache, Keep the Beat for Real-Time Interaction Model Serving

语音交互 | 8.7/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6175 字 阅读 →
论文解读

SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models

语音交互 | 8.9/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7382 字 阅读 →
论文解读

\(\tau\)-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains

语音交互 | 9.1/10

 · 更新于 2026-09-24 · 约 18 分钟 · 8721 字 阅读 →
论文解读

BFCL Audio: An Audio Function Calling Evaluation for Large Language Models

语音交互 | 7.7/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7260 字 阅读 →
论文解读

LALM-as-a-Judge: Benchmarking Large Audio-Language Models for Safety Evaluation in Multi-Turn Spoken Dialogues

语音交互 | 8.1/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7266 字 阅读 →