论文解读

Automatic Stage Lighting Control: Is it a Rule-Driven Process or Generative Task?

音乐生成 | 8.5/10

 · 更新于 2026-09-11 · 约 11 分钟 · 5357 字 阅读 →
论文解读

AVERE: Improving Audiovisual Emotion Reasoning with Preference Optimization

情感识别 | 7.5/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4542 字 阅读 →
论文解读

AVEX: What Matters for Animal Vocalization Encoding

生物声学 | 7.5/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4577 字 阅读 →
论文解读

AVoCaDO: An Audiovisual Video Captioner Driven by Temporal Orchestration

音视频 | 8.0/10

 · 更新于 2026-09-11 · 约 13 分钟 · 6017 字 阅读 →
论文解读

Better Together: Leveraging Unpaired Multimodal Data for Stronger Unimodal Models

多模态模型 | 7.0/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4906 字 阅读 →
论文解读

Beyond Instance-Level Alignment: Dual-Level Optimal Transport for Audio-Text Retrieval

音频检索 | 8.0/10

 · 更新于 2026-09-11 · 约 9 分钟 · 4293 字 阅读 →
论文解读

Bridging Piano Transcription and Rendering via Disentangled Score Content and Style

音乐信息检索 | 8.0/10

 · 更新于 2026-09-11 · 约 11 分钟 · 5197 字 阅读 →
论文解读

Can Speech LLMs Think while Listening?

语音对话系统 | 7.5/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4627 字 阅读 →
论文解读

Can Vision-Language Models Answer Face to Face Questions in the Real-World?

音频问答 | 7.5/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4693 字 阅读 →
论文解读

Characterizing and Optimizing the Spatial Kernel of Multi Resolution Hash Encodings

3D重建 | 7.0/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4516 字 阅读 →
论文解读

Closing the Gap Between Text and Speech Understanding in LLMs

语音对话系统 | 7.5/10

 · 更新于 2026-09-11 · 约 17 分钟 · 8119 字 阅读 →
论文解读

Compose and Fuse: Revisiting the Foundational Bottlenecks in Multimodal Reasoning

多模态推理 | 8.5/10

 · 更新于 2026-09-11 · 约 8 分钟 · 3952 字 阅读 →
论文解读

Confident and Adaptive Generative Speech Recognition via Risk Control

语音识别 | 8.0/10

 · 更新于 2026-09-11 · 约 11 分钟 · 5063 字 阅读 →
论文解读

Continuous Audio Language Models

音频生成 音乐生成 | 9.5/10

 · 更新于 2026-09-11 · 约 11 分钟 · 5470 字 阅读 →
论文解读

CTC-DRO: Robust Optimization for Reducing Language Disparities in Speech Recognition

语音识别 | 8.0/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4981 字 阅读 →
论文解读

Data-Centric Lessons To Improve Speech-Language Pretraining

语音问答 | 8.0/10

 · 更新于 2026-09-11 · 约 8 分钟 · 3893 字 阅读 →
论文解读

Deep Learning with Learnable Product-Structured Activations

音频分类 | 7.5/10

 · 更新于 2026-09-11 · 约 11 分钟 · 5033 字 阅读 →
论文解读

DiffSDA: Unsupervised Diffusion Sequential Disentanglement Across Modalities

无监督学习 | 8.0/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4931 字 阅读 →
论文解读

Discovering and Steering Interpretable Concepts in Large Generative Music Models

音乐生成 | 7.5/10

 · 更新于 2026-09-11 · 约 9 分钟 · 4300 字 阅读 →
论文解读

DiVeQ: Differentiable Vector Quantization Using the Reparameterization Trick

语音编码 | 8.5/10

 · 更新于 2026-09-11 · 约 13 分钟 · 6275 字 阅读 →