论文解读

Fast When, Careful Who: Dual-Process Multiparty Turn-Taking with Diffusion Augmentation

语音活动检测 | 5.9/10

 · 更新于 2026-09-08 · 约 13 分钟 · 6103 字 阅读 →
论文解读

FreeSonic: Training-Free Temporal-Aware Decoupled Attention for Precise Audio Editing

音频生成 | 9/10

 · 更新于 2026-09-08 · 约 12 分钟 · 5891 字 阅读 →
论文解读

From Awareness to Adherence: Bridging the Context Gap in Spoken Dialogue Systems via Context-Aware Decoding

语音识别 | 6.7/10

 · 更新于 2026-09-08 · 约 5 分钟 · 2144 字 阅读 →
论文解读

From Physics to Representation: Audio Learning with Synthetic Pre-training via Procedural Generation

自监督学习 | 8.2/10

 · 更新于 2026-09-08 · 约 11 分钟 · 5339 字 阅读 →
论文解读

Geometrically Constrained Decentralized Independent Vector Analysis for Distributed Microphone Arrays

语音分离 | 7.2/10

 · 更新于 2026-09-08 · 约 12 分钟 · 5760 字 阅读 →
论文解读

Interpretable and Frugal Learning Systems Employing Multiresolution Pyramids and Volterra Kernels

Interpretable and Frugal Learning Systems Employing Multiresolution Pyramids and Volterra Kernels

 · 更新于 2026-09-08 · 约 11 分钟 · 5374 字 阅读 →
论文解读

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction

语音合成 | 6.8/10

 · 更新于 2026-09-08 · 约 11 分钟 · 5195 字 阅读 →
论文解读

Learning Input-Channel Permutation Equivariance for Multi-Channel Source Separation: Reducing Bleeding in Small Music Ensembles

音乐源分离 | 7.9/10

 · 更新于 2026-09-08 · 约 11 分钟 · 5210 字 阅读 →
论文解读

LLM-Based Synthetic Ground Truth Generation for Audio-Based Emotion Classification via In-Context Learning

数据增强 | 5.3/10

 · 更新于 2026-09-08 · 约 12 分钟 · 5545 字 阅读 →
论文解读

MAF: Multimodal Adaptive Few-shot Prompting for Sentiment Analysis with MLLMs

多模态模型 | 5.9/10

 · 更新于 2026-09-08 · 约 13 分钟 · 6343 字 阅读 →
论文解读

MambAdapter: Lightweight Mamba-Based Adapters for Parameter-Efficient Transfer Learning in Speech and Audio

语音识别 | 8.9/10

 · 更新于 2026-09-08 · 约 10 分钟 · 4872 字 阅读 →
论文解读

MatchLM2Lite: A Scalable MLLM-to-Lite Framework for Reproduced Content Identification

音频分类 | 8.3/10

 · 更新于 2026-09-08 · 约 13 分钟 · 6080 字 阅读 →
论文解读

MUNI: Multimodal Unified Latent Diffusion for Coherent Any-to-Any Generation

语音生成 | 6.9/10

 · 更新于 2026-09-08 · 约 11 分钟 · 5373 字 阅读 →
论文解读

MuVAP: Multimodal Multiparty Voice Activity Projection for Turn-taking Prediction in the Wild

语音对话系统 | 7.8/10

 · 更新于 2026-09-08 · 约 14 分钟 · 6606 字 阅读 →
论文解读

NVMOS: Non-Verbal Vocalization Quality Assessment in Speech

自监督学习 | 6.2/10

 · 更新于 2026-09-08 · 约 10 分钟 · 4781 字 阅读 →
论文解读

Phonetically Explainable Speech Deepfake Detection

语音伪造检测 | 9/10

 · 更新于 2026-09-08 · 约 15 分钟 · 7309 字 阅读 →
论文解读

Pixel-TTS: Image based Text Rendering for Robust Text-to-Speech

语音合成 | 7.5/10

 · 更新于 2026-09-08 · 约 12 分钟 · 5784 字 阅读 →
论文解读

Probing Low Frame Rate Degradation in Neural Audio Codecs

语音生成 | 8.6/10

 · 更新于 2026-09-08 · 约 11 分钟 · 5018 字 阅读 →
论文解读

Rhythm of the Deep: A Computational-Linguistic Test of Duality of Patterning in Sperm Whale Codas

自监督学习 | 8.5/10

 · 更新于 2026-09-08 · 约 14 分钟 · 6634 字 阅读 →
论文解读

Robust Spoofed Speech Detection via Temporal Pyramid Modeling

音频深度伪造检测 | 6.7/10

 · 更新于 2026-09-08 · 约 14 分钟 · 6597 字 阅读 →