论文解读

Towards Understanding Modality Interaction in Multimodal Language Models via Partial Information Decomposition

Towards Understanding Modality Interaction in Multimodal Language Models via Partial Information Decomposition

 · 更新于 2026-09-09 · 约 1 分钟 · 38 字 阅读 →
论文解读

Two-dimensional quantization for geometry-aware audio coding

Two-dimensional quantization for geometry-aware audio coding

 · 更新于 2026-09-09 · 约 1 分钟 · 32 字 阅读 →
论文解读

Unlocking Speech–Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning

Unlocking Speech–Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning

 · 更新于 2026-09-09 · 约 1 分钟 · 47 字 阅读 →
论文解读

Verifiable Multimodal Reasoning: Fact-level Attribution with Multimodal Sources

Verifiable Multimodal Reasoning: Fact-level Attribution with Multimodal Sources

 · 更新于 2026-09-09 · 约 1 分钟 · 34 字 阅读 →
论文解读

VIBE: Disentangling Social Dynamics via Kinematics-Informed Variational Inference for Behavioral Emotion

VIBE: Disentangling Social Dynamics via Kinematics-Informed Variational Inference for Behavioral Emotion

 · 更新于 2026-09-09 · 约 1 分钟 · 37 字 阅读 →
论文解读

video-SALMONN S: Memory-Enhanced Streaming Audio-Visual LLM

video-SALMONN S: Memory-Enhanced Streaming Audio-Visual LLM

 · 更新于 2026-09-09 · 约 1 分钟 · 32 字 阅读 →
论文解读

VocSim A Training-free Benchmark for Zero-shot Content Identity in Single-source Audio

VocSim A Training-free Benchmark for Zero-shot Content Identity in Single-source Audio

 · 更新于 2026-09-09 · 约 1 分钟 · 37 字 阅读 →
论文解读

WaveSSM: Multiscale State-Space Models for Non-stationary Signal Attention

WaveSSM: Multiscale State-Space Models for Non-stationary Signal Attention

 · 更新于 2026-09-09 · 约 1 分钟 · 34 字 阅读 →
论文解读

Zero-Shot Rankability: Revealing Latent Ordinal Structure in Multimodal Large Language Models via Language

Zero-Shot Rankability: Revealing Latent Ordinal Structure in Multimodal Large Language Models via Language

 · 更新于 2026-09-09 · 约 1 分钟 · 39 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-23

共分析 123 篇语音/AI 论文

 · 更新于 2026-09-09 · 约 10 分钟 · 4688 字 阅读 →
论文解读

Academic Text-to-Music Grand Challenge: Datasets, Baselines, and Evaluation Methods

音乐生成 | 9.9/10

 · 更新于 2026-09-09 · 约 15 分钟 · 7170 字 阅读 →
论文解读

Automatic Contextual Audio Denoising

语音去噪 | 7.5/10

 · 更新于 2026-09-09 · 约 11 分钟 · 5457 字 阅读 →
论文解读

Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models

语音情感识别 | 7.0/10

 · 更新于 2026-09-09 · 约 15 分钟 · 7101 字 阅读 →
论文解读

Do Factual Recall Mechanisms Carry over from Text to Speech in Multimodal Language Models?

大语言模型 | 10.0/10

 · 更新于 2026-09-09 · 约 10 分钟 · 4998 字 阅读 →
论文解读

Effective User-defined Keyword Spotting with Dual-stage Matching, Multi-modal Enrollment, and Continual Adaptation

关键词检测 | 7.4/10

 · 更新于 2026-09-09 · 约 13 分钟 · 6327 字 阅读 →
论文解读

From Volterra Series to Kunchenko Stochastic Polynomials: Half a Century of Non-Gaussian Estimation Methodology

统计信号处理 | 7.8/10

 · 更新于 2026-09-09 · 约 12 分钟 · 5756 字 阅读 →
论文解读

In Silico Modeling of the RAMPHO Buffer: Dissociating Informational and Energetic Masking via Phonetic Entropy in Deep Neural Networks

认知科学 | 6.5/10

 · 更新于 2026-09-09 · 约 11 分钟 · 5140 字 阅读 →
论文解读

LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning

跨模态 | 9.0/10

 · 更新于 2026-09-09 · 约 12 分钟 · 5861 字 阅读 →
论文解读

Live Music Diffusion Models: Efficient Fine-Tuning and Post-Training of Interactive Diffusion Music Generators

音乐生成 | 5.9/10

 · 更新于 2026-09-09 · 约 16 分钟 · 7711 字 阅读 →
论文解读

MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue

跨模态 | 6.5/10

 · 更新于 2026-09-09 · 约 13 分钟 · 6048 字 阅读 →