论文解读

Speech Entrainment in Multi-Party Conversations with a Digital Agent

语音交互 | 5.3/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5599 字 阅读 →
论文解读

Speech Signals Complement LLMs for Predicting Interpersonal Attraction in Speed Dating

语音交互 | 6.9/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7083 字 阅读 →
论文解读

StanceBench: A Benchmark for Audio LLM-Based Interpersonal Stance Evaluation from Speech

语音属性识别 | 8.6/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8263 字 阅读 →
论文解读

Towards High-Level Semantic Intelligence

音视频理解 | 7.0/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6680 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-07-28

共分析 31 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 103 分钟 · 51404 字 阅读 →
论文解读

CODA: Cascaded Online Discontinuity-Aware Alignment for Real-Time Image-Based Score Following

音频理解 | 8.1/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6774 字 阅读 →
论文解读

How Meta-Learning Shapes LoRA Adapter Geometry in Speech Deepfake Detection

语音伪造检测 | 5.2/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7098 字 阅读 →
论文解读

Kutti AI: A Voice-First, Offline-Capable Learning Companion with Real-Time Struggle Detection for Visually-Impaired Children

语音交互 | 5.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4866 字 阅读 →
论文解读

Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding

语音识别 | 6.6/10

 · 更新于 2026-09-06 · 约 19 分钟 · 9351 字 阅读 →
论文解读

MemNMF: Memory-Augmented NMF on LPC Spectra for Anomalous Sound Detection

音频事件检测 | 5.3/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6236 字 阅读 →
论文解读

MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond

语音识别 | 8.2/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6955 字 阅读 →
论文解读

Music-JEPA: Learning a World Model of Sound from Action

音乐转录 | 6.2/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6393 字 阅读 →
论文解读

Phylogenetic signal in marine mammal and bird vocalizations captured by audio foundation models: the limited benefit of domain-specific pretraining

音频理解 | 7.7/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6153 字 阅读 →
论文解读

Probing Speaker Identity Sensitivity in Audio Deepfake Detectors

语音伪造检测 | 6.5/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6060 字 阅读 →
论文解读

Reflector: Arrangement-Aware Harmonic Retrieval for Sample-Based Composition

音乐检索 | 7.7/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5708 字 阅读 →
论文解读

SoundscapeAgent: Agentic Soundscape Construction for Controllable Synthesis and Scalable Audio-Language Supervision

大语言模型 | 8.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5047 字 阅读 →
论文解读

Synthetic Speech, Real Signal: Paralinguistic Preservation and Cross-Lingual Augmentation via Voice Cloning

语音情感识别 | 6.2/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6667 字 阅读 →
论文解读

Transforming Keystroke Noise to Text: Self-Supervised Acoustic Eavesdropping Attacks on Keyboards

语音活动检测 | 6.5/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7374 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-07-27

共分析 13 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 39 分钟 · 19087 字 阅读 →
论文解读

An Evaluation Framework for Structured Audio Captions Validated by Controlled Perturbations

数据集 | 5.3/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7447 字 阅读 →