论文解读

Beyond Mapping: Domain-Invariant Representations via Spectral Embedding of Optimal Transport Plans

领域适应 | 7.5/10

 · 更新于 2026-09-14 · 约 12 分钟 · 5892 字 阅读 →
论文解读

Bimodal Fusion Framework for Dynamic Facial Expression Recognition In-The-Wild

语音情感识别 | 7.0/10

 · 更新于 2026-09-14 · 约 8 分钟 · 3875 字 阅读 →
论文解读

BioSEN: A Bio-Acoustic Signal Enhancement Network for Animal Vocalizations

生物声学 | 7.5/10

 · 更新于 2026-09-14 · 约 11 分钟 · 5221 字 阅读 →
论文解读

BiRQ: Bi-Level Self-Labeling Random Quantization for Self-Supervised Speech Recognition

语音识别 | 8.0/10

 · 更新于 2026-09-14 · 约 10 分钟 · 4898 字 阅读 →
论文解读

Bleed No More: Generative Interference Reduction for Musical Recordings

音乐源分离 | 7.0/10

 · 更新于 2026-09-14 · 约 10 分钟 · 4891 字 阅读 →
论文解读

Bloodroot: When Watermarking Turns Poisonous for Stealthy Backdoor

音频安全 | 7.5/10

 · 更新于 2026-09-14 · 约 10 分钟 · 4790 字 阅读 →
论文解读

Bone-Conduction Guided Multimodal Speech Enhancement with Conditional Diffusion Models

语音增强 | 7.5/10

 · 更新于 2026-09-14 · 约 9 分钟 · 4225 字 阅读 →
论文解读

Brainprint-Modulated Target Speaker Extraction

语音分离 | 8.0/10

 · 更新于 2026-09-14 · 约 10 分钟 · 4574 字 阅读 →
论文解读

Break-the-Beat! Controllable MIDI-to-Drum audio synthesis

音乐生成 | 7.5/10

 · 更新于 2026-09-14 · 约 9 分钟 · 4419 字 阅读 →
论文解读

BridgeCode: A Dual Speech Representation Paradigm for Autoregressive Zero-Shot Text-to-Speech Synthesis

语音合成 | 8.0/10

 · 更新于 2026-09-14 · 约 11 分钟 · 5196 字 阅读 →
论文解读

Bridging the Front-End and Back-End for Robust ASR via Cross-Attention-Based U-Net

语音识别 | 7.0/10

 · 更新于 2026-09-14 · 约 10 分钟 · 4857 字 阅读 →
论文解读

Bridging the Measurement–Simulation Gap in Room Acoustics with Real2sim Diffusion

声源定位 | 8.5/10

 · 更新于 2026-09-14 · 约 9 分钟 · 4409 字 阅读 →
论文解读

Bridging the Semantic Gap: Cross-Attentive Fusion for Joint Acoustic-Semantic Speech Quality Assessment

语音质量评估 | 8.5/10

 · 更新于 2026-09-14 · 约 12 分钟 · 5819 字 阅读 →
论文解读

BSMP-SENet:Band-Split Magnitude-Phase Network for Speech Enhancement

语音增强 | 7.0/10

 · 更新于 2026-09-14 · 约 9 分钟 · 4436 字 阅读 →
论文解读

CALM: Joint Contextual Acoustic-Linguistic Modeling for Personalization of Multi-Speaker ASR

语音识别 | 7.5/10

 · 更新于 2026-09-14 · 约 11 分钟 · 5139 字 阅读 →
论文解读

CaMoD: Causal-Aware Modality Denoising for Multimodal Dialogue Intent Recognition

多模态对话意图识别 | 7.5/10

 · 更新于 2026-09-14 · 约 10 分钟 · 4555 字 阅读 →
论文解读

Can Hierarchical Cross-Modal Fusion Predict Human Perception of AI Dubbed Content?

模型评估 | 6.0/10

 · 更新于 2026-09-14 · 约 9 分钟 · 4504 字 阅读 →
论文解读

Can Large Audio Language Models Understand Audio Well? Speech, Scene and Events Understanding Benchmark for LALMs

基准测试 | 7.0/10

 · 更新于 2026-09-14 · 约 10 分钟 · 4576 字 阅读 →
论文解读

Caption and Audio-Guided Video Representation Learning with Gated Attention for Partially Relevant Video Retrieval

视频检索 | 7.0/10

 · 更新于 2026-09-14 · 约 9 分钟 · 4469 字 阅读 →
论文解读

Cardiobridge-DM: Bridging Cross-Cohort Heart Sound Synthesis via Rhythm-Aware Semi-Supervised Diffusion

音频生成 | 7.5/10

 · 更新于 2026-09-14 · 约 10 分钟 · 4823 字 阅读 →