论文解读

Full-Duplex Interaction in Spoken Dialogue Systems: A Comprehensive Study from the ICASSP 2026 HumDial Challenge

语音对话系统 | 6.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3578 字 阅读 →
论文解读

MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction

语音对话系统 | 8.5/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6664 字 阅读 →
论文解读

Predicting Upcoming Stuttering Events from Three-Second Audio: Stratified Evaluation Reveals Severity-Selective Precursors, and the Model Deploys Fully On-Device

语音生物标志物 | 7.0/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6353 字 阅读 →
论文解读

Text-Utilization for Encoder-dominated Speech Recognition Models

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 6 分钟 · 2856 字 阅读 →
论文解读

A Speech-Driven Paradigm for Physics-Informed Modeling of Coupled Micro-Speakers

音频生成 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4668 字 阅读 →
论文解读

Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5056 字 阅读 →
论文解读

Advancing LLM-Based Multi-Channel Multi-Speaker Speech Recognition with Global Cross-Channel Attention and Sentence-Ordered First-In First-Out Serialized Output Training

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5058 字 阅读 →
论文解读

ALMA-Chor: Leveraging Audio-Lyric Alignment with Mamba for Chorus Detection

音乐信息检索 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4379 字 阅读 →
论文解读

An End-to-End Multimodal System for Subtitle Recognition and Chinese-Japanese Translation in Short Dramas

多模态模型 | 7.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5044 字 阅读 →
论文解读

An Envelope Separation Aided Multi-Task Learning Model for Blind Source Counting and Localization

声源定位 | 6.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4186 字 阅读 →
论文解读

Audio Deepfake Detection at the First Greeting: "Hi!"

音频深度伪造检测 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4507 字 阅读 →
论文解读

Audio-to-Score Jazz Solo Transcription with the Rhythm Perceiver

音乐信息检索 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4493 字 阅读 →
论文解读

Auditory-Inspired Transformer for Binaural Speech Enhancement and Spatial Cue Preservation

语音增强 | 7.0/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5802 字 阅读 →
论文解读

CALM: Joint Contextual Acoustic-Linguistic Modeling for Personalization of Multi-Speaker ASR

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5139 字 阅读 →
论文解读

Chunk-Wise Attention Transducers for Fast and Accurate Streaming Speech-to-Text

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4722 字 阅读 →
论文解读

Chunkwise Aligners for Streaming Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4405 字 阅读 →
论文解读

Content Anonymization for Privacy in Long-Form Audio

语音匿名化 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4691 字 阅读 →
论文解读

Deep Dubbing: End-to-End Auto-Audiobook System with Text-to-Timbre and Context-Aware Instruct-TTS

语音合成 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4986 字 阅读 →
论文解读

Direct Transfer of Prosody in Speech-to-speech Translation using Disentangled Speech Tokens

语音翻译 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4896 字 阅读 →
论文解读

Discrete-Continuous Fusion With Adaptive Hierarchical Features For Audio Deepfake Detection

音频深度伪造检测 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4332 字 阅读 →