论文解读

AUV: Teaching Audio Universal Vector Quantization with Single Nested Codebook

音频生成 | 8.0/10

 · 更新于 2026-09-14 · 约 10 分钟 · 4987 字 阅读 →
论文解读

Auxiliary Multi-Label Training For Improving the Robustness of Audio Deepfake Detection on AI-Processed Data

音频深度伪造检测 | 6.5/10

 · 更新于 2026-09-14 · 约 8 分钟 · 3885 字 阅读 →
论文解读

AVATAR: Audio-Visual Adaptive Fusion via Trained Agent Reinforcement for Multimodal Deepfake Detection

音频深度伪造检测 | 7.5/10

 · 更新于 2026-09-14 · 约 9 分钟 · 4343 字 阅读 →
论文解读

AVO-65: A Large-Scale Hierarchical Audio-Visual Object Dataset

音视频 | 7.0/10

 · 更新于 2026-09-14 · 约 9 分钟 · 4186 字 阅读 →
论文解读

B-GRPO: Unsupervised Speech Emotion Recognition Based on Batched-Group Relative Policy Optimization

语音情感识别 | 6.5/10

 · 更新于 2026-09-14 · 约 9 分钟 · 4316 字 阅读 →
论文解读

BACHI: Boundary-Aware Symbolic Chord Recognition Through Masked Iterative Decoding on POP and Classical Music

音乐信息检索 | 7.5/10

 · 更新于 2026-09-14 · 约 9 分钟 · 4376 字 阅读 →
论文解读

Bayesian Low-Rank Factorization for Robust Model Adaptation

语音识别 | 8.0/10

 · 更新于 2026-09-14 · 约 9 分钟 · 4176 字 阅读 →
论文解读

Bayesian Signal Separation Via Plug-and-Play Diffusion-Within-Gibbs Sampling

语音分离 | 7.5/10

 · 更新于 2026-09-14 · 约 10 分钟 · 4956 字 阅读 →
论文解读

BBPE16: UTF-16-Based Byte-Level Byte-Pair Encoding for Improved Multilingual Speech Recognition

语音识别 | 7.0/10

 · 更新于 2026-09-14 · 约 7 分钟 · 3372 字 阅读 →
论文解读

Beamforming Using Virtual Microphones for Hearing Aid Applications

语音增强 | 7.5/10

 · 更新于 2026-09-14 · 约 8 分钟 · 3819 字 阅读 →
论文解读

Beat and Downbeat Detection: A Reformulated Approach

音乐理解 | 7.5/10

 · 更新于 2026-09-14 · 约 9 分钟 · 4272 字 阅读 →
论文解读

BeatMamba: Bidirectional Selective State-Space Modeling for Efficient Beat Tracking

音乐信息检索 | 7.5/10

 · 更新于 2026-09-14 · 约 10 分钟 · 4564 字 阅读 →
论文解读

Behind the Scenes: Mechanistic Interpretability of Lora-Adapted Whisper for Speech Emotion Recognition

语音情感识别 | 7.5/10

 · 更新于 2026-09-14 · 约 8 分钟 · 3951 字 阅读 →
论文解读

Benchmarking Humans And Machines On Complex Multilingual Speech Understanding Tasks

音频问答 | 7.5/10

 · 更新于 2026-09-14 · 约 8 分钟 · 3566 字 阅读 →
论文解读

Benchmarking Music Autotagging with MGPHot Expert Annotations vs. Generic Tag Datasets

音乐信息检索 | 7.5/10

 · 更新于 2026-09-14 · 约 9 分钟 · 4442 字 阅读 →
论文解读

BEST-RQ-based Self-Supervised Learning for Whisper Domain Adaptation

语音识别 | 7.5/10

 · 更新于 2026-09-14 · 约 11 分钟 · 5130 字 阅读 →
论文解读

BEST-STD 2.0: Balanced and Efficient Speech Tokenizer for Spoken Term Detection

音频检索 | 7.5/10

 · 更新于 2026-09-14 · 约 11 分钟 · 5339 字 阅读 →
论文解读

Beyond Face Swapping: A Diffusion-Based Digital Human Benchmark for Multimodal Deepfake Detection

音频深度伪造检测 | 8.1/10

 · 更新于 2026-09-14 · 约 10 分钟 · 4573 字 阅读 →
论文解读

Beyond Global Emotion: Fine-Grained Emotional Speech Synthesis with Dynamic Word-Level Modulation

语音合成 | 7.5/10

 · 更新于 2026-09-14 · 约 10 分钟 · 4555 字 阅读 →
论文解读

Beyond Isolated Utterances: Cue-Guided Interaction for Context-Dependent Conversational Multimodal Understanding

多模态模型 | 7.5/10

 · 更新于 2026-09-14 · 约 12 分钟 · 5721 字 阅读 →