论文解读

AVERE: Improving Audiovisual Emotion Reasoning with Preference Optimization

语音情感识别 | 8.0/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4850 字 阅读 →
论文解读

AVEX: What Matters for Animal Vocalization Encoding

生物声学 | 7.0/10

 · 更新于 2026-09-10 · 约 11 分钟 · 5485 字 阅读 →
论文解读

AVoCaDO: An Audiovisual Video Captioner Driven by Temporal Orchestration

视频描述生成 | 8.5/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4718 字 阅读 →
论文解读

Better Together: Leveraging Unpaired Multimodal Data for Stronger Unimodal Models

音频分类 | 7.0/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4990 字 阅读 →
论文解读

Beyond Decodability: Reconstructing Language Model Representations with an Encoding Probe

模型评估 | 7.5/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4999 字 阅读 →
论文解读

Beyond Instance-Level Alignment: Dual-Level Optimal Transport for Audio-Text Retrieval

音频检索 | 7.5/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4524 字 阅读 →
论文解读

Bridging Piano Transcription and Rendering via Disentangled Score Content and Style

音乐信息检索 | 8.0/10

 · 更新于 2026-09-10 · 约 13 分钟 · 6223 字 阅读 →
论文解读

Can Speech LLMs Think while Listening?

语音对话系统 | 7.5/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4703 字 阅读 →
论文解读

Can Vision-Language Models Answer Face to Face Questions in the Real-World?

音频问答 | 8.0/10

 · 更新于 2026-09-10 · 约 9 分钟 · 4402 字 阅读 →
论文解读

Closing the Gap Between Text and Speech Understanding in LLMs

语音大模型 | 8.5/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4518 字 阅读 →
论文解读

Compose and Fuse: Revisiting the Foundational Bottlenecks in Multimodal Reasoning

多模态推理 | 7.5/10

 · 更新于 2026-09-10 · 约 9 分钟 · 4438 字 阅读 →
论文解读

Confident and Adaptive Generative Speech Recognition via Risk Control

语音识别 | 7.5/10

 · 更新于 2026-09-10 · 约 9 分钟 · 4401 字 阅读 →
论文解读

Continuous Audio Language Models

语音合成 | 7.0/10

 · 更新于 2026-09-10 · 约 12 分钟 · 5685 字 阅读 →
论文解读

CTC-DRO: Robust Optimization for Reducing Language Disparities in Speech Recognition

语音识别 | 9.0/10

 · 更新于 2026-09-10 · 约 9 分钟 · 4230 字 阅读 →
论文解读

CustomDancer: Customized Dance Recommendation by Text-Dance Retrieval

音频检索 音乐理解 | 6.5/10

 · 更新于 2026-09-10 · 约 12 分钟 · 5689 字 阅读 →
论文解读

Data-Centric Lessons To Improve Speech-Language Pretraining

语音问答 | 8.0/10

 · 更新于 2026-09-10 · 约 9 分钟 · 4090 字 阅读 →
论文解读

Deep Learning with Learnable Product-Structured Activations

神经网络架构 | 8.0/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4694 字 阅读 →
论文解读

DiffSDA: Unsupervised Diffusion Sequential Disentanglement Across Modalities

序列解耦 | 8.0/10

 · 更新于 2026-09-10 · 约 13 分钟 · 6401 字 阅读 →
论文解读

Discovering and Steering Interpretable Concepts in Large Generative Music Models

音乐生成 | 8.0/10

 · 更新于 2026-09-10 · 约 9 分钟 · 4454 字 阅读 →
论文解读

DiVeQ: Differentiable Vector Quantization Using the Reparameterization Trick

生成模型 | 8.0/10

 · 更新于 2026-09-10 · 约 12 分钟 · 5814 字 阅读 →