论文解读

Multi-Modal Semantic Expansion with Constrained LLM Reranking for Conversational Music Recommendation

音乐推荐 | 8.3/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4665 字 阅读 →
论文解读

TurboBias 2.0: Streaming Context-Biasing for Production-Efficient ASR Systems

语音识别 | 9.0/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5444 字 阅读 →
论文解读

Multimodal Rapport Estimation in Real-World HRI

多模态模型 | 7.0/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4533 字 阅读 →
论文解读

AT-ADD: All-Type Audio Deepfake Detection Challenge Summary

音频伪造检测 | 6.3/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7256 字 阅读 →
论文解读

Smartphone Audio Based Distress Detection

音频事件检测 | 6.3/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7186 字 阅读 →
论文解读

Music-Source-Separation-Training (MSST): A Unified Framework for Training and Evaluating Music Demixing Models

音乐源分离 | 5.7/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7537 字 阅读 →
论文解读

AlphaWiSE: Adaptive Weight Interpolation for Continual Multimodal Representation Learning

音频检索 | 6.4/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7023 字 阅读 →
论文解读

Unlocking Speech–Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning

语音交互 | 6.7/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7653 字 阅读 →
论文解读

Unlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning

语音交互 | 8.5/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5306 字 阅读 →
论文解读

MuVAP: Multimodal Multiparty Voice Activity Projection for Turn-taking Prediction in the Wild

语音对话系统 | 7.8/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6606 字 阅读 →
论文解读

Stabilizing Short Duration Speaker Verification through Neural Re-scoring with Hybrid Enrollment

说话人验证 | 7.9/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5781 字 阅读 →
论文解读

Conan-embedding-v3: Fusing Modality-Specific Models for Omni-Modal Embedding

音频检索 | 7.6/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5579 字 阅读 →
论文解读

Parameter-efficient Dual-encoder Architecture with Differentiable Choquet Integral Fusion for Underwater Acoustic Classification

音频分类 | 6.4/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5654 字 阅读 →
论文解读

Thaka at KSAA-2026 Task 2: Regularized Fine-Tuning for Arabic Speech Diacritization

语音增强 | 5.5/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5587 字 阅读 →
论文解读

Diffusion Domain Expansion: Learning to Coordinate Pre-trained Diffusion Models

扩散模型 | 7.4/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5697 字 阅读 →
论文解读

Robust and Lightweight F0 Estimation Through Mid-Level Fusion of DSP-Informed Features

基频估计 | 8.0/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5195 字 阅读 →