论文解读

Stable Spectral Copula Alignment for Robust Multimodal Learning

鲁棒性 | 5.2/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7374 字 阅读 →
论文解读

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits

音视频生成 | 6.8/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7540 字 阅读 →
论文解读

T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation

音视频生成 | 7.9/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8106 字 阅读 →
论文解读

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions

音视频理解 | 9.4/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7829 字 阅读 →
论文解读

Towards Understanding Modality Interaction in Multimodal Language Models via Partial Information Decomposition

音视频理解 | 5.3/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6064 字 阅读 →
论文解读

V-LynX: Token Interface Alignment for Video+X LLMs

音视频问答 | 7.8/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5123 字 阅读 →
论文解读

video-SALMONN S: Memory-Enhanced Streaming Audio-Visual LLM

音视频问答 | 7.3/10

 · 更新于 2026-09-06 · 约 20 分钟 · 9623 字 阅读 →
论文解读

Zero-Shot Rankability: Revealing Latent Ordinal Structure in Multimodal Large Language Models via Language

音视频理解 | 6.8/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6730 字 阅读 →
论文解读

A global predicted-fMRI drive signal from TRIBE does not predict YouTube replay heatmaps

音视频理解 | 7.7/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7749 字 阅读 →
论文解读

An Efficient vLLM-Based Inference Pipeline for Unified Audio Understanding and Generation

语音合成 | 6.8/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6302 字 阅读 →
论文解读

Reasoning LLM Improves Speaker Recognition in Long-form TV Dramas

音视频理解 | 7.2/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7157 字 阅读 →
论文解读

ASR-Agnostic Multimodal Spectrotemporal Modeling for Early Dementia Detection

多模态模型 | 7.4/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8321 字 阅读 →
论文解读

AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation

语音合成 | 6.5/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6668 字 阅读 →
论文解读

Beyond Binary Instrument QA: Probing Instrument Grounding in Music Audio-Language Models

音频分类 | 7.6/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4548 字 阅读 →
论文解读

Building a Multimodal Dataset of Academic Paper for Keyword Extraction

语音识别 | 5.2/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5542 字 阅读 →
论文解读

Listening Between the Lines: Joint Learning of ASR Embeddings and LLM-Augmented Linguistics for Dementia Detection

数据增强 | 7/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6468 字 阅读 →
论文解读

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs

语音合成 | 7.9/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6024 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-07-01

共分析 35 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 102 分钟 · 50963 字 阅读 →
论文解读

AMR: Adaptive Modality Routing for Multimodal Polyglot Speaker Identification

说话人识别 | 7.8/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6161 字 阅读 →
论文解读

Effective Depth in Joint Source-Channel Coding: An Implicit Equilibrium Analysis

语音编码 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4439 字 阅读 →