论文解读

FlexiCodec: A Dynamic Neural Audio Codec for Low Frame Rates

语音合成 | 8.8/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6115 字 阅读 →
论文解读

MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction

语音对话系统 | 8.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5987 字 阅读 →
论文解读

STITCH: Simultaneous Thinking and Talking with Chunked Reasoning for Spoken Language Models

语音对话系统 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3665 字 阅读 →
论文解读

Towards True Speech-to-Speech Models Without Text Guidance

语音对话系统 | 9.1/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4831 字 阅读 →
论文解读

TVTSyn: Content-Synchronous Time-Varying Timbre for Streaming Voice Conversion and Anonymization

语音转换 | 7.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6178 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-02

共分析 4 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 15 分钟 · 7026 字 阅读 →
论文解读

MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction

语音对话系统 | 8.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6664 字 阅读 →
论文解读

Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation

音视频 | 8.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5319 字 阅读 →
论文解读

Text-Utilization for Encoder-dominated Speech Recognition Models

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 6 分钟 · 2856 字 阅读 →
论文解读

A Generative-First Neural Audio Autoencoder

音乐生成 | 8.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5163 字 阅读 →
论文解读

An Efficient Neural Network for Modeling Human Auditory Neurograms for Speech

语音增强 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4940 字 阅读 →
论文解读

Chunk-Wise Attention Transducers for Fast and Accurate Streaming Speech-to-Text

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4722 字 阅读 →
论文解读

Chunkwise Aligners for Streaming Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4405 字 阅读 →
论文解读

CTC-DID: CTC-Based Arabic Dialect Identification for Streaming Applications

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4764 字 阅读 →
论文解读

Direct Simultaneous Translation Activation for Large Audio-Language Models

语音翻译 | 6.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4476 字 阅读 →
论文解读

Do we really need self-attention for streaming automatic speech recognition?

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4519 字 阅读 →
论文解读

EEND-SAA: Enrollment-Less Main Speaker Voice Activity Detection Using Self-Attention Attractors

语音活动检测 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4559 字 阅读 →
论文解读

Entropy-Guided GRVQ for Ultra-Low Bitrate Neural Speech Codec

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4068 字 阅读 →
论文解读

Equipping Large Language Model with Directional Speech Understanding Capabilities

语音识别 语音翻译 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4467 字 阅读 →
论文解读

FastEnhancer: Speed-Optimized Streaming Neural Speech Enhancement

语音增强 | 8.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5043 字 阅读 →