论文解读

A Cross-lingual Comparison of Human and Classification Model Entrainment Behavior in Code-switched Speech Settings

Transformer | 5.1/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7784 字 阅读 →
论文解读

AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities

多模态模型 | 8.3/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5386 字 阅读 →
论文解读

Depression Markers in Speech: An Approach based on Tract Variables Dynamics

语音属性识别 | 5.1/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6197 字 阅读 →
论文解读

Device Invariance using Domain Adaptation on Acoustic Scene Classification

领域适应 | 4.2/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6363 字 阅读 →
论文解读

DynaBridge: Dynamic Summary-Guided Cross-Task Multimodal Fusion for DASS-Structured Mental Health Assessment

音视频理解 | 4.9/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8175 字 阅读 →
论文解读

Evaluation of forced alignment of code-mixed speech: the case of Hindi-English

语音识别 | 5.4/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4432 字 阅读 →
论文解读

Finding the noise: Zero-shot AI Music Detection

无监督学习 | 5.1/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6037 字 阅读 →
论文解读

From Semantics to Readout: Mechanistic Understanding of Audio Tokens after Fine-Tuning for Temporal Audio Grounding

音频理解 | 5.1/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6416 字 阅读 →
论文解读

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models

音视频问答 | 8.3/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8316 字 阅读 →
论文解读

Spacing Out: On the Reliability of Binaural Music Source Separation Metrics

音乐源分离 | 6.1/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6032 字 阅读 →
论文解读

SpeechLLM Meets Federated Learning for End-to-End ASR: English and Italian Case Studies

语音识别 | 7.1/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5520 字 阅读 →
论文解读

VAD to the Bone: Ultra-Tiny Speech Activity Detection for Edge Deployment

语音活动检测 | 7.4/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6296 字 阅读 →
论文解读

Disentangling Acoustic Cues in Alzheimer's Pathology and Perception: The Roles of Language and Gender

语音属性识别 | 5.4/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7657 字 阅读 →
论文解读

Do Visual Features Improve Other-Initiated Repair Detection? A Dyadic Multimodal Approach

音视频交互 | 5.9/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7206 字 阅读 →
论文解读

Explainable AI through the Lens of Material Agency: Enabling Musical Interface Design with Neural Audio Models

音乐生成 | 6.7/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6159 字 阅读 →
论文解读

Expose Your Disguise: Recovering Source Speaker Identity From Voice Conversion

对比学习 | 7.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6629 字 阅读 →
论文解读

Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages

说话人日志 | 7.1/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8448 字 阅读 →
论文解读

Infinite Canons: Maximally Self-Similar Melodic Lines and Canons with Infinite Solutions

音乐生成 | 6.4/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6656 字 阅读 →
论文解读

Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers

语音合成 | 6.8/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7089 字 阅读 →
论文解读

Mind the Microphone Gap: Benchmarking Array Upsampling Strategies for Latent Acoustic Mapping

声源定位 | 6.6/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6635 字 阅读 →