论文解读

Voice Biomarkers for Depression and Anxiety

语音生物标志物 | 1.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4782 字 阅读 →
论文解读

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM

语音大模型 | 8.0/10

 · 更新于 2026-09-25 · 约 29 分钟 · 14118 字 阅读 →
论文解读

MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model

语音对话系统 | 7.5/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7885 字 阅读 →
论文解读

A Comprehensive Analysis of Tokenization and Self-Supervised Learning in End-to-End Automatic Speech Recognition applied on French Language

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4807 字 阅读 →
论文解读

MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model

语音对话系统 | 7.5/10

 · 更新于 2026-09-25 · 约 26 分钟 · 12846 字 阅读 →
论文解读

A cross-species neural foundation model for end-to-end speech decoding

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5724 字 阅读 →
论文解读

Automatic Stage Lighting Control: Is it a Rule-Driven Process or Generative Task?

音乐生成 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4931 字 阅读 →
论文解读

End-to-end Listen, Look, Speak and Act

语音对话系统 | 8.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5323 字 阅读 →
论文解读

From Text to Talk: Audio-Language Model Needs Non-Autoregressive Joint Training

语音对话系统 | 8.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5679 字 阅读 →
论文解读

Hierarchical Semantic-Acoustic Modeling via Semi-Discrete Residual Representations for Expressive End-to-End Speech Synthesis

语音合成 | 8.0/10

 · 更新于 2026-09-25 · 约 20 分钟 · 9658 字 阅读 →
论文解读

LadderSym: A Multimodal Interleaved Transformer for Music Practice Error Detection

音乐理解 | 8.0/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5557 字 阅读 →
论文解读

Learnable Fractional Superlets with a Spectro-Temporal Emotion Encoder for Speech Emotion Recognition

语音情感识别 | 8.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5399 字 阅读 →
论文解读

RoboOmni: Proactive Robot Manipulation in Omni-modal Context

机器人操作 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5562 字 阅读 →
论文解读

STITCH: Simultaneous Thinking and Talking with Chunked Reasoning for Spoken Language Models

语音对话系统 | 8.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4500 字 阅读 →
论文解读

TINY BUT MIGHTY: A SOFTWARE-HARDWARE CO- DESIGN APPROACH FOR EFFICIENT MULTIMODAL IN- FERENCE ON BATTERY-POWERED SMALL DEVICES

多模态模型 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4229 字 阅读 →
论文解读

Towards True Speech-to-Speech Models Without Text Guidance

语音对话系统 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5080 字 阅读 →
论文解读

TripleSumm: Adaptive Triple-Modality Fusion for Video Summarization

视频摘要 | 8.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5060 字 阅读 →
论文解读

UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice

语音翻译 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4434 字 阅读 →
论文解读

A cross-species neural foundation model for end-to-end speech decoding

语音识别 | 8.0/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5641 字 阅读 →
论文解读

Automatic Stage Lighting Control: Is it a Rule-Driven Process or Generative Task?

音乐生成 | 8.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5357 字 阅读 →