论文解读

The Role of Disfluencies in Speech Translation

语音翻译 | 6.4/10

 · 更新于 2026-09-25 · 约 22 分钟 · 10618 字 阅读 →
论文解读

Uncertainty-Aware Crossmodal Fusion for Classification of Animal Behavior

音频分类 | 6.3/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7297 字 阅读 →
论文解读

A Neurosymbolic Approach for Explainable Early Diagnosis of Alzheimer's Disease

音频理解 | 7.7/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7773 字 阅读 →
论文解读

Cloned Voices, Real Consequences: Evaluating Bias in Political Deepfake Detection for Electoral Integrity in Brazil

语音伪造检测 | 7.7/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7303 字 阅读 →
论文解读

Do Music Foundation Models Embed Pitch in Helical Structure?

音乐理解 | 8.1/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6831 字 阅读 →
论文解读

DoubleHelix: Structured Cross-Modal Fusion for Audio-Visual Speech Recognition with LLMs

音视频语音识别 | 7.4/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9385 字 阅读 →
论文解读

Exploring Efficient Waveform Diffusion Models for Foley Sound Generation

音频生成 | 6.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6173 字 阅读 →
论文解读

FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models

音视频理解 | 7.9/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7441 字 阅读 →
论文解读

Learning to Predict Performance-induced Emotion Differences in Classical Piano Music

数据集 | 5.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5773 字 阅读 →
论文解读

Leveraging Beam Search Information for Confidence Estimation in E2E ASR

语音识别 | 6.3/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7632 字 阅读 →
论文解读

M3-DuplexBench: A Multi-Turn, Multilingual, Multidomain Benchmark for Full-Duplex Spoken Dialogue Models

语音交互 | 6.1/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5540 字 阅读 →
论文解读

Model-Agnostic Meta-Learning Initialization for Distributed Multichannel Active Noise Control

主动降噪 | 5.1/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5740 字 阅读 →
论文解读

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens

语音合成 | 5.8/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7230 字 阅读 →
论文解读

Technological Advances in Detecting and Managing Cognitive Impairment in Older Adults: Trends, Challenges, and Future Directions

医疗音频 | 4.0/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6862 字 阅读 →
论文解读

Tensor-Based Joint Pitch and DOA Estimation

声源定位 | 6.9/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6290 字 阅读 →
论文解读

TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models

音频理解 | 6.1/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8405 字 阅读 →
论文解读

Versatile On-device Adaptation at the Edge by Unifying Few-shot, Zero-shot, Continual, and In-context Learning

元学习 | 8.1/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7787 字 阅读 →
论文解读

ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine

音视频理解 | 8.3/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5406 字 阅读 →
论文解读

AgenticASR: Refining Speech Recognition in Real-World Scenarios via an Agentic Approach

语音识别 | 6.8/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6977 字 阅读 →
论文解读

Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPO

语音交互 | 7.2/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8276 字 阅读 →