论文解读

Brainprint-Modulated Target Speaker Extraction

语音分离 | 8.0/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4574 字 阅读 →
论文解读

Break-the-Beat! Controllable MIDI-to-Drum audio synthesis

音乐生成 | 7.5/10

 · 更新于 2026-09-11 · 约 9 分钟 · 4419 字 阅读 →
论文解读

BridgeCode: A Dual Speech Representation Paradigm for Autoregressive Zero-Shot Text-to-Speech Synthesis

语音合成 | 8.0/10

 · 更新于 2026-09-11 · 约 11 分钟 · 5196 字 阅读 →
论文解读

Bridging the Front-End and Back-End for Robust ASR via Cross-Attention-Based U-Net

语音识别 | 7.0/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4857 字 阅读 →
论文解读

Bridging the Measurement–Simulation Gap in Room Acoustics with Real2sim Diffusion

声源定位 | 8.5/10

 · 更新于 2026-09-11 · 约 9 分钟 · 4409 字 阅读 →
论文解读

Bridging the Semantic Gap: Cross-Attentive Fusion for Joint Acoustic-Semantic Speech Quality Assessment

语音质量评估 | 8.5/10

 · 更新于 2026-09-11 · 约 12 分钟 · 5819 字 阅读 →
论文解读

BSMP-SENet:Band-Split Magnitude-Phase Network for Speech Enhancement

语音增强 | 7.0/10

 · 更新于 2026-09-11 · 约 9 分钟 · 4436 字 阅读 →
论文解读

CALM: Joint Contextual Acoustic-Linguistic Modeling for Personalization of Multi-Speaker ASR

语音识别 | 7.5/10

 · 更新于 2026-09-11 · 约 11 分钟 · 5139 字 阅读 →
论文解读

CaMoD: Causal-Aware Modality Denoising for Multimodal Dialogue Intent Recognition

多模态对话意图识别 | 7.5/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4555 字 阅读 →
论文解读

Can Hierarchical Cross-Modal Fusion Predict Human Perception of AI Dubbed Content?

模型评估 | 6.0/10

 · 更新于 2026-09-11 · 约 9 分钟 · 4504 字 阅读 →
论文解读

Can Large Audio Language Models Understand Audio Well? Speech, Scene and Events Understanding Benchmark for LALMs

基准测试 | 7.0/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4576 字 阅读 →
论文解读

Caption and Audio-Guided Video Representation Learning with Gated Attention for Partially Relevant Video Retrieval

视频检索 | 7.0/10

 · 更新于 2026-09-11 · 约 9 分钟 · 4469 字 阅读 →
论文解读

Cardiobridge-DM: Bridging Cross-Cohort Heart Sound Synthesis via Rhythm-Aware Semi-Supervised Diffusion

音频生成 | 7.5/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4823 字 阅读 →
论文解读

CASTELLA: Long Audio Dataset with Captions and Temporal Boundaries

音频检索 | 8.5/10

 · 更新于 2026-09-11 · 约 8 分钟 · 3821 字 阅读 →
论文解读

CCST: Cross-Modal and Consistency-Aware Self-Training for Source-Free Unsupervised Domain Adaptation in Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-11 · 约 13 分钟 · 6342 字 阅读 →
论文解读

Chunk-Wise Attention Transducers for Fast and Accurate Streaming Speech-to-Text

语音识别 | 7.5/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4722 字 阅读 →
论文解读

Chunkwise Aligners for Streaming Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-11 · 约 9 分钟 · 4405 字 阅读 →
论文解读

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources

音频场景理解 | 7.5/10

 · 更新于 2026-09-11 · 约 9 分钟 · 4177 字 阅读 →
论文解读

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents

基准测试 | 7.0/10

 · 更新于 2026-09-11 · 约 8 分钟 · 3824 字 阅读 →
论文解读

Clue2Emo: A Brain-Inspired Framework for Open-Vocabulary Multimodal Emotion Recognition

语音情感识别 | 8.5/10

 · 更新于 2026-09-11 · 约 12 分钟 · 6003 字 阅读 →