论文解读

VIBE: Disentangling Social Dynamics via Kinematics-Informed Variational Inference for Behavioral Emotion

VIBE: Disentangling Social Dynamics via Kinematics-Informed Variational Inference for Behavioral Emotion

 · 更新于 2026-10-01 · 约 1 分钟 · 37 字 阅读 →
论文解读

video-SALMONN S: Memory-Enhanced Streaming Audio-Visual LLM

video-SALMONN S: Memory-Enhanced Streaming Audio-Visual LLM

 · 更新于 2026-10-01 · 约 1 分钟 · 32 字 阅读 →
论文解读

VocSim A Training-free Benchmark for Zero-shot Content Identity in Single-source Audio

VocSim A Training-free Benchmark for Zero-shot Content Identity in Single-source Audio

 · 更新于 2026-10-01 · 约 1 分钟 · 37 字 阅读 →
论文解读

WaveSSM: Multiscale State-Space Models for Non-stationary Signal Attention

WaveSSM: Multiscale State-Space Models for Non-stationary Signal Attention

 · 更新于 2026-10-01 · 约 1 分钟 · 34 字 阅读 →
论文解读

Zero-Shot Rankability: Revealing Latent Ordinal Structure in Multimodal Large Language Models via Language

Zero-Shot Rankability: Revealing Latent Ordinal Structure in Multimodal Large Language Models via Language

 · 更新于 2026-10-01 · 约 1 分钟 · 39 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-23

共分析 123 篇语音/AI 论文

 · 更新于 2026-10-01 · 约 10 分钟 · 4688 字 阅读 →
论文解读

Academic Text-to-Music Grand Challenge: Datasets, Baselines, and Evaluation Methods

音乐生成 | 9.9/10

 · 更新于 2026-10-01 · 约 15 分钟 · 7170 字 阅读 →
论文解读

Automatic Contextual Audio Denoising

语音去噪 | 7.5/10

 · 更新于 2026-10-01 · 约 11 分钟 · 5457 字 阅读 →
论文解读

Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models

语音情感识别 | 7.0/10

 · 更新于 2026-10-01 · 约 15 分钟 · 7101 字 阅读 →
论文解读

Do Factual Recall Mechanisms Carry over from Text to Speech in Multimodal Language Models?

大语言模型 | 10.0/10

 · 更新于 2026-10-01 · 约 10 分钟 · 4998 字 阅读 →
论文解读

Effective User-defined Keyword Spotting with Dual-stage Matching, Multi-modal Enrollment, and Continual Adaptation

关键词检测 | 7.4/10

 · 更新于 2026-10-01 · 约 13 分钟 · 6327 字 阅读 →
论文解读

From Volterra Series to Kunchenko Stochastic Polynomials: Half a Century of Non-Gaussian Estimation Methodology

统计信号处理 | 7.8/10

 · 更新于 2026-10-01 · 约 12 分钟 · 5756 字 阅读 →
论文解读

In Silico Modeling of the RAMPHO Buffer: Dissociating Informational and Energetic Masking via Phonetic Entropy in Deep Neural Networks

认知科学 | 6.5/10

 · 更新于 2026-10-01 · 约 11 分钟 · 5140 字 阅读 →
论文解读

LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning

跨模态 | 9.0/10

 · 更新于 2026-10-01 · 约 12 分钟 · 5861 字 阅读 →
论文解读

Live Music Diffusion Models: Efficient Fine-Tuning and Post-Training of Interactive Diffusion Music Generators

音乐生成 | 5.9/10

 · 更新于 2026-10-01 · 约 16 分钟 · 7711 字 阅读 →
论文解读

MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue

跨模态 | 6.5/10

 · 更新于 2026-10-01 · 约 13 分钟 · 6048 字 阅读 →
论文解读

Neighbor-Consistent Neural Filters for Robust Personal Sound Zones Under Localization Uncertainty

声区控制 | 8.5/10

 · 更新于 2026-10-01 · 约 18 分钟 · 8836 字 阅读 →
论文解读

OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding

音视频 | 7.3/10

 · 更新于 2026-10-01 · 约 13 分钟 · 6324 字 阅读 →
论文解读

Plug-in Losses for Evidential Deep Learning: A Simplified Framework for Uncertainty Estimation that Includes the Softmax Classifier

模型评估 | 3.5/10

 · 更新于 2026-10-01 · 约 13 分钟 · 6504 字 阅读 →
论文解读

Real-time, EDM-inspired sonfication of the activity of a supercomputer

数据声化 | 6.5/10

 · 更新于 2026-10-01 · 约 12 分钟 · 5745 字 阅读 →