论文解读

DBFT-SD: Weakly Supervised Multimodal Detection of Sensitive Audio-Visual Content

音频事件检测 | 8.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 4008 字 阅读 →
论文解读

Distilling Attention Knowledge for Speaker Verification

说话人验证 | 8.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3961 字 阅读 →
论文解读

EchoRAG: A Two-Stage Framework for Audio-Text Retrieval and Temporal Grounding

音频检索 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5031 字 阅读 →
论文解读

EdgeSpot: Efficient and High-Performance Few-Shot Model for Keyword Spotting

语音活动检测 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4076 字 阅读 →
论文解读

Enabling Multi-Species Bird Classification on Low-Power Bioacoustic Loggers

生物声学 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4664 字 阅读 →
论文解读

Enhancing Speaker Verification with w2v-BERT 2.0 and Knowledge Distillation Guided Structured Pruning

说话人验证 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4903 字 阅读 →
论文解读

FocalCodec-Stream: Streaming Low-Bitrate Speech Coding via Causal Distillation

语音编码 | 8.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5230 字 阅读 →
论文解读

From Hallucination to Articulation: Language Model-Driven Losses for Ultra Low-Bitrate Neural Speech Coding

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4405 字 阅读 →
论文解读

GLUE: Gradient-free Learning to Unify Experts

迁移学习 | 6.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5139 字 阅读 →
论文解读

Int-MeanFlow: Few-Step Speech Generation with Integral Velocity Distillation

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5048 字 阅读 →
论文解读

Learning to Align with Unbalanced Optimal Transport in Linguistic Knowledge Transfer for ASR

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4273 字 阅读 →
论文解读

Lightweight and Generalizable Acoustic Scene Representations Via Contrastive Fine-Tuning and Distillation

音频场景理解 | 8.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5226 字 阅读 →
论文解读

MI-Fuse: Label Fusion for Unsupervised Domain Adaptation with Closed-Source Large Audio-Language Model

语音情感识别 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4330 字 阅读 →
论文解读

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation

音频生成 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5908 字 阅读 →
论文解读

Prompt-Guided Mixture-of-Experts for Robust Multimodal Sentiment Analysis with Missing Modalities

语音情感识别 | 8.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4723 字 阅读 →
论文解读

S-SONDO: Self-Supervised Knowledge Distillation for General Audio Foundation Models

音频分类 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4416 字 阅读 →
论文解读

Salad-VAE: Semantic Audio Compression with Language-Audio Distillation

音频压缩 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4645 字 阅读 →
论文解读

Semantic Anchor Transfer from Short to Long Speech in a Distillation-Based Summarization Framework

语音摘要 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5537 字 阅读 →
论文解读

SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models

音频问答 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5286 字 阅读 →
论文解读

Sounds that Shape: Audio-Driven 3D Mesh Generation with Attribute-Decoupled Score Distillation Sampling

音频生成 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3847 字 阅读 →