论文解读

Bayesian Signal Separation Via Plug-and-Play Diffusion-Within-Gibbs Sampling

语音分离 | 7.5/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4956 字 阅读 →
论文解读

BBPE16: UTF-16-Based Byte-Level Byte-Pair Encoding for Improved Multilingual Speech Recognition

语音识别 | 7.0/10

 · 更新于 2026-09-11 · 约 7 分钟 · 3372 字 阅读 →
论文解读

Beamforming Using Virtual Microphones for Hearing Aid Applications

语音增强 | 7.5/10

 · 更新于 2026-09-11 · 约 8 分钟 · 3819 字 阅读 →
论文解读

Beat and Downbeat Detection: A Reformulated Approach

音乐理解 | 7.5/10

 · 更新于 2026-09-11 · 约 9 分钟 · 4272 字 阅读 →
论文解读

BeatMamba: Bidirectional Selective State-Space Modeling for Efficient Beat Tracking

音乐信息检索 | 7.5/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4564 字 阅读 →
论文解读

Behind the Scenes: Mechanistic Interpretability of Lora-Adapted Whisper for Speech Emotion Recognition

语音情感识别 | 7.5/10

 · 更新于 2026-09-11 · 约 8 分钟 · 3951 字 阅读 →
论文解读

Benchmarking Humans And Machines On Complex Multilingual Speech Understanding Tasks

音频问答 | 7.5/10

 · 更新于 2026-09-11 · 约 8 分钟 · 3566 字 阅读 →
论文解读

Benchmarking Music Autotagging with MGPHot Expert Annotations vs. Generic Tag Datasets

音乐信息检索 | 7.5/10

 · 更新于 2026-09-11 · 约 9 分钟 · 4442 字 阅读 →
论文解读

BEST-RQ-based Self-Supervised Learning for Whisper Domain Adaptation

语音识别 | 7.5/10

 · 更新于 2026-09-11 · 约 11 分钟 · 5130 字 阅读 →
论文解读

BEST-STD 2.0: Balanced and Efficient Speech Tokenizer for Spoken Term Detection

音频检索 | 7.5/10

 · 更新于 2026-09-11 · 约 11 分钟 · 5339 字 阅读 →
论文解读

Beyond Face Swapping: A Diffusion-Based Digital Human Benchmark for Multimodal Deepfake Detection

音频深度伪造检测 | 8.1/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4573 字 阅读 →
论文解读

Beyond Global Emotion: Fine-Grained Emotional Speech Synthesis with Dynamic Word-Level Modulation

语音合成 | 7.5/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4555 字 阅读 →
论文解读

Beyond Isolated Utterances: Cue-Guided Interaction for Context-Dependent Conversational Multimodal Understanding

多模态模型 | 7.5/10

 · 更新于 2026-09-11 · 约 12 分钟 · 5721 字 阅读 →
论文解读

Beyond Mapping: Domain-Invariant Representations via Spectral Embedding of Optimal Transport Plans

领域适应 | 7.5/10

 · 更新于 2026-09-11 · 约 12 分钟 · 5892 字 阅读 →
论文解读

Bimodal Fusion Framework for Dynamic Facial Expression Recognition In-The-Wild

语音情感识别 | 7.0/10

 · 更新于 2026-09-11 · 约 8 分钟 · 3875 字 阅读 →
论文解读

BioSEN: A Bio-Acoustic Signal Enhancement Network for Animal Vocalizations

生物声学 | 7.5/10

 · 更新于 2026-09-11 · 约 11 分钟 · 5221 字 阅读 →
论文解读

BiRQ: Bi-Level Self-Labeling Random Quantization for Self-Supervised Speech Recognition

语音识别 | 8.0/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4898 字 阅读 →
论文解读

Bleed No More: Generative Interference Reduction for Musical Recordings

音乐源分离 | 7.0/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4891 字 阅读 →
论文解读

Bloodroot: When Watermarking Turns Poisonous for Stealthy Backdoor

音频安全 | 7.5/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4790 字 阅读 →
论文解读

Bone-Conduction Guided Multimodal Speech Enhancement with Conditional Diffusion Models

语音增强 | 7.5/10

 · 更新于 2026-09-11 · 约 9 分钟 · 4225 字 阅读 →