Towards Understanding Modality Interaction in Multimodal Language Models via Partial Information Decomposition
Towards Understanding Modality Interaction in Multimodal Language Models via Partial Information Decomposition
Towards Understanding Modality Interaction in Multimodal Language Models via Partial Information Decomposition
Two-dimensional quantization for geometry-aware audio coding
Unlocking Speech–Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning
Verifiable Multimodal Reasoning: Fact-level Attribution with Multimodal Sources
VIBE: Disentangling Social Dynamics via Kinematics-Informed Variational Inference for Behavioral Emotion
video-SALMONN S: Memory-Enhanced Streaming Audio-Visual LLM
VocSim A Training-free Benchmark for Zero-shot Content Identity in Single-source Audio
WaveSSM: Multiscale State-Space Models for Non-stationary Signal Attention
Zero-Shot Rankability: Revealing Latent Ordinal Structure in Multimodal Large Language Models via Language
共分析 123 篇语音/AI 论文
音乐生成 | 9.9/10
语音去噪 | 7.5/10
语音情感识别 | 7.0/10
大语言模型 | 10.0/10
关键词检测 | 7.4/10
统计信号处理 | 7.8/10
认知科学 | 6.5/10
跨模态 | 9.0/10
音乐生成 | 5.9/10
跨模态 | 6.5/10