论文解读

How Reliable Are Multimodal Signals of Conversational State? Evidence from Remote Dyadic Collaborative Tasks

鲁棒性 | 6.6/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7971 字 阅读 →
论文解读

Is One Score Enough? Assessing Singing Quality of Songs with Temporal Score Curves

音乐理解 | 7.6/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7120 字 阅读 →
论文解读

NABEATs: Noise-Aware Audio Representation Learning

音频理解 | 6.7/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8604 字 阅读 →
论文解读

Should Missing Modalities Always Be Necessary to Repair for Multi-modal Sentiment Analysis?

多模态模型 | 7.0/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7627 字 阅读 →
论文解读

A Study of Parallelizable Alternatives to Dynamic Time Warping for Aligning Long Sequences

基准测试 | 8.1/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6393 字 阅读 →
论文解读

AnovaX: A Local, Multi-Agent Voice Assistant with LLM Planning, Typed Executors, and Adaptive Recovery

语音交互 | 4.8/10

 · 更新于 2026-09-25 · 约 21 分钟 · 10217 字 阅读 →
论文解读

Data-driven Video Codec with Implicit Neural Representations

音频编码 | 5.3/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6797 字 阅读 →
论文解读

Estimating the Reliability of Dynamic Time Warping Alignments Using Circumstantial Evidence

音乐理解 | 7.6/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6737 字 阅读 →
论文解读

Proof-Carrying Multimodal Timelines: Finite-Trace Modal Certificates for Video-Audio Consistency

基准测试 | 8.6/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8117 字 阅读 →
论文解读

Segmental DTW: A Parallelizable Alternative to Dynamic Time Warping

音频检索 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5278 字 阅读 →
论文解读

StemFX: Learning Mixing Style Representations via Autoregressive FX Chain Prediction on Source-Separated Stems

自回归模型 | 9.6/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8092 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-07-20

共分析 15 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 52 分钟 · 25690 字 阅读 →
论文解读

MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation

音视频生成 | 6.3/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6689 字 阅读 →
论文解读

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems

语音合成 | 7.9/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9074 字 阅读 →
论文解读

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment

多模态模型 | 5.6/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6202 字 阅读 →
论文解读

Video = World + Event Stream

音频理解 | 4.9/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6888 字 阅读 →
论文解读

What does the model actually see? Evaluation protocols and input availability in data-driven prediction of room acoustic parameters

音频质量评估 | 7.2/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7819 字 阅读 →
论文解读

Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation

语音质量评估 | 8.2/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7473 字 阅读 →
论文解读

Cover First, Disagree Softly: Rethinking Mismatch-First Active Learning for Frame-Level Audio Classification

音频事件检测 | 6.7/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5869 字 阅读 →
论文解读

Genre Bias or Aesthetic Perception? Identifying and Mitigating Shortcut Learning in Music Evaluation

音乐理解 | 6.0/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6179 字 阅读 →