论文解读

A Study of Parallelizable Alternatives to Dynamic Time Warping for Aligning Long Sequences

基准测试 | 8.1/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6393 字 阅读 →
论文解读

AnovaX: A Local, Multi-Agent Voice Assistant with LLM Planning, Typed Executors, and Adaptive Recovery

语音交互 | 4.8/10

 · 更新于 2026-09-06 · 约 21 分钟 · 10217 字 阅读 →
论文解读

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos

音视频理解 | 6.0/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7850 字 阅读 →
论文解读

AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis

语音合成 | 6.9/10

 · 更新于 2026-09-06 · 约 18 分钟 · 8526 字 阅读 →
论文解读

AV-JEPA: Extending LeJEPA to Audio-Visual Self-Supervised Learning

音视频理解 | 5.7/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7410 字 阅读 →
论文解读

Constrained Hebbian Learning Supports Efficient Representational Allocation under Structural Constraints

音视频理解 | 6.7/10

 · 更新于 2026-09-06 · 约 29 分钟 · 14475 字 阅读 →
论文解读

Controlling Implicit Shortcut Reliance in L2 Spoken English Auto-markers

语音质量评估 | 7.5/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8422 字 阅读 →
论文解读

Data-driven Video Codec with Implicit Neural Representations

音频编码 | 5.3/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6797 字 阅读 →
论文解读

Estimating the Reliability of Dynamic Time Warping Alignments Using Circumstantial Evidence

音乐理解 | 7.6/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6737 字 阅读 →
论文解读

Natural Backdoor Attacks on Speech Recognition Models

语音识别 | 3.5/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7287 字 阅读 →
论文解读

Proof-Carrying Multimodal Timelines: Finite-Trace Modal Certificates for Video-Audio Consistency

基准测试 | 8.6/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8117 字 阅读 →
论文解读

Segmental DTW: A Parallelizable Alternative to Dynamic Time Warping

音频检索 | 7.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5278 字 阅读 →
论文解读

SpeechGuard: Online Defense against Backdoor Attacks on Speech Recognition Models

语音识别 | 6.0/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6373 字 阅读 →
论文解读

StemFX: Learning Mixing Style Representations via Autoregressive FX Chain Prediction on Source-Separated Stems

自回归模型 | 9.6/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8092 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-07-20

共分析 15 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 52 分钟 · 25690 字 阅读 →
论文解读

AlphaWiSE: Adaptive Weight Interpolation for Continual Multimodal Representation Learning

音频检索 | 6.4/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7023 字 阅读 →
论文解读

Can Tokens Compete? Token Representations against Supervised CNN Backbones for BirdCLEF+ 2026

音频事件检测 | 8.3/10

 · 更新于 2026-09-06 · 约 19 分钟 · 9493 字 阅读 →
论文解读

Dialogs: a studio-quality expressive conversational Russian speech corpus for dialog assistants

语音合成 | 7.8/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6565 字 阅读 →
论文解读

InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring

多模态模型 | 7.3/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5751 字 阅读 →
论文解读

ITGPT: A Transformer Based Architecture for the Generation of Dance Dance Revolution and In the Groove Charts

音乐生成 | 6.5/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7824 字 阅读 →