论文解读

MFCL Audio: An Audio Function Calling Evaluation for Large Language Models

MFCL Audio: An Audio Function Calling Evaluation for Large Language Models

 · 更新于 2026-09-09 · 约 1 分钟 · 37 字 阅读 →
论文解读

MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models

MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models

 · 更新于 2026-09-09 · 约 1 分钟 · 35 字 阅读 →
论文解读

MoST: Mixing Speech and Text with Modality-Aware Mixture of Experts

MoST: Mixing Speech and Text with Modality-Aware Mixture of Experts

 · 更新于 2026-09-09 · 约 1 分钟 · 36 字 阅读 →
论文解读

MOTOR: A Multimodal Dataset for Two-Wheeler Rider Behavior Understanding

视频行为识别 | 5.9/10

 · 更新于 2026-09-09 · 约 11 分钟 · 5468 字 阅读 →
论文解读

Multimodal Fusion via Self-Consistent Task-Gradient Fields

Multimodal Fusion via Self-Consistent Task-Gradient Fields

 · 更新于 2026-09-09 · 约 1 分钟 · 32 字 阅读 →
论文解读

Multimodal Latent Language Modeling with Next-Token Diffusion

Multimodal Latent Language Modeling with Next-Token Diffusion

 · 更新于 2026-09-09 · 约 1 分钟 · 33 字 阅读 →
论文解读

Multiple Choice Learning of Low-Rank Adapters for Language Modeling

Multiple Choice Learning of Low-Rank Adapters for Language Modeling

 · 更新于 2026-09-09 · 约 1 分钟 · 35 字 阅读 →
论文解读

MusicDET: Zero-Shot AI-Generated Music Detection

MusicDET: Zero-Shot AI-Generated Music Detection

 · 更新于 2026-09-09 · 约 1 分钟 · 31 字 阅读 →
论文解读

NAACA: Training-Free NeuroAuditory Attentive Cognitive Architecture with Oscillatory Working Memory for Salience-Driven Attention Gating

NAACA: Training-Free NeuroAuditory Attentive Cognitive Architecture with Oscillatory Working Memory for Salience-Driven Attention Gating

 · 更新于 2026-09-09 · 约 1 分钟 · 40 字 阅读 →
论文解读

Neural-Inspired Modeling of Auditory Selection and Compensation for Audio-Visual Speech Separation

Neural-Inspired Modeling of Auditory Selection and Compensation for Audio-Visual Speech Separation

 · 更新于 2026-09-09 · 约 1 分钟 · 37 字 阅读 →
论文解读

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning

 · 更新于 2026-09-09 · 约 1 分钟 · 33 字 阅读 →
论文解读

OmniDenseCap: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions

OmniDenseCap: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions

 · 更新于 2026-09-09 · 约 1 分钟 · 36 字 阅读 →
论文解读

OmniShow: Orchestrating Multimodal Conditions for Human-Object Interaction Video Generation

OmniShow: Orchestrating Multimodal Conditions for Human-Object Interaction Video Generation

 · 更新于 2026-09-09 · 约 1 分钟 · 35 字 阅读 →
论文解读

OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models

OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models

 · 更新于 2026-09-09 · 约 1 分钟 · 36 字 阅读 →
论文解读

OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention

OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention

 · 更新于 2026-09-09 · 约 1 分钟 · 36 字 阅读 →
论文解读

Optimality of FSQ tokens for continuous diffusion for categorical data with application to text-to-speech

Optimality of FSQ tokens for continuous diffusion for categorical data with application to text-to-speech

 · 更新于 2026-09-09 · 约 1 分钟 · 40 字 阅读 →
论文解读

PADS-TAL: Padding-Annealed Diffusion Sampling in Text-Aware Latent Space for Robust and Diverse Text-to-Music Generation

PADS-TAL: Padding-Annealed Diffusion Sampling in Text-Aware Latent Space for Robust and Diverse Text-to-Music Generation

 · 更新于 2026-09-09 · 约 1 分钟 · 40 字 阅读 →
论文解读

PCRNet: Phase-aware Complex Refinement Network for EEG-based Auditory Attention Decoding

PCRNet: Phase-aware Complex Refinement Network for EEG-based Auditory Attention Decoding

 · 更新于 2026-09-09 · 约 1 分钟 · 36 字 阅读 →
论文解读

PHALAR: Phasors for Learned Musical Audio Representations

PHALAR: Phasors for Learned Musical Audio Representations

 · 更新于 2026-09-09 · 约 1 分钟 · 33 字 阅读 →
论文解读

PhaseCoder: Microphone Geometry-Agnostic Spatial Audio Understanding for Multimodal LLMs

PhaseCoder: Microphone Geometry-Agnostic Spatial Audio Understanding for Multimodal LLMs

 · 更新于 2026-09-09 · 约 1 分钟 · 35 字 阅读 →