论文解读

Advancing LLM-Based Multi-Channel Multi-Speaker Speech Recognition with Global Cross-Channel Attention and Sentence-Ordered First-In First-Out Serialized Output Training

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5058 字 阅读 →
论文解读

ALMA-Chor: Leveraging Audio-Lyric Alignment with Mamba for Chorus Detection

音乐信息检索 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4379 字 阅读 →
论文解读

An End-to-End Multimodal System for Subtitle Recognition and Chinese-Japanese Translation in Short Dramas

多模态模型 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5044 字 阅读 →
论文解读

An Envelope Separation Aided Multi-Task Learning Model for Blind Source Counting and Localization

声源定位 | 6.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4186 字 阅读 →
论文解读

Audio Deepfake Detection at the First Greeting: "Hi!"

音频深度伪造检测 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4507 字 阅读 →
论文解读

Audio-to-Score Jazz Solo Transcription with the Rhythm Perceiver

音乐信息检索 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4493 字 阅读 →
论文解读

Auditory-Inspired Transformer for Binaural Speech Enhancement and Spatial Cue Preservation

语音增强 | 7.0/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5802 字 阅读 →
论文解读

CALM: Joint Contextual Acoustic-Linguistic Modeling for Personalization of Multi-Speaker ASR

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5139 字 阅读 →
论文解读

Chunk-Wise Attention Transducers for Fast and Accurate Streaming Speech-to-Text

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4722 字 阅读 →
论文解读

Chunkwise Aligners for Streaming Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4405 字 阅读 →
论文解读

Content Anonymization for Privacy in Long-Form Audio

语音匿名化 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4691 字 阅读 →
论文解读

Deep Dubbing: End-to-End Auto-Audiobook System with Text-to-Timbre and Context-Aware Instruct-TTS

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4986 字 阅读 →
论文解读

Direct Transfer of Prosody in Speech-to-speech Translation using Disentangled Speech Tokens

语音翻译 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4896 字 阅读 →
论文解读

Discrete-Continuous Fusion With Adaptive Hierarchical Features For Audio Deepfake Detection

音频深度伪造检测 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4332 字 阅读 →
论文解读

DSRMS-TransUnet: A Decentralized Non-Shifted Transunet for Shallow Water Acoustic Source Range Estimation

声源定位 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4613 字 阅读 →
论文解读

Dual-Strategy-Enhanced Conbimamba for Neural Speaker Diarization

说话人分离 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4642 字 阅读 →
论文解读

E2E-AEC: Implementing An End-To-End Neural Network Learning Approach for Acoustic Echo Cancellation

语音增强 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4163 字 阅读 →
论文解读

EEND-SAA: Enrollment-Less Main Speaker Voice Activity Detection Using Self-Attention Attractors

语音活动检测 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4559 字 阅读 →
论文解读

Exploring SSL Discrete Tokens for Multilingual Automatic Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5356 字 阅读 →
论文解读

HAVT-IVD: Heterogeneity-Aware Cross-Modal Network for Audio-Visual Surveillance: Idling Vehicles Detection with Multichannel Audio and Multiscale Visual Cues

音频事件检测 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4420 字 阅读 →