论文解读

BAT: Better Audio Transformer Guided by Convex Gated Probing

BAT: Better Audio Transformer Guided by Convex Gated Probing

 · 更新于 2026-09-09 · 约 1 分钟 · 35 字 阅读 →
论文解读

BEAT: Tokenizing and Generating Symbolic Music by Uniform Temporal Steps

BEAT: Tokenizing and Generating Symbolic Music by Uniform Temporal Steps

 · 更新于 2026-09-09 · 约 1 分钟 · 36 字 阅读 →
论文解读

Bioacoustic Geolocation: Species Sounds as Geographic Signals

Bioacoustic Geolocation: Species Sounds as Geographic Signals

 · 更新于 2026-09-09 · 约 1 分钟 · 33 字 阅读 →
论文解读

Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models

Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models

 · 更新于 2026-09-09 · 约 1 分钟 · 41 字 阅读 →
论文解读

Bridging Your Imagination with Audio-Video Generation via a Unified Director

Bridging Your Imagination with Audio-Video Generation via a Unified Director

 · 更新于 2026-09-09 · 约 1 分钟 · 36 字 阅读 →
论文解读

Characterizing the Predictive Impact of Modalities with Supervised Latent-Variable Modeling

Characterizing the Predictive Impact of Modalities with Supervised Latent-Variable Modeling

 · 更新于 2026-09-09 · 约 1 分钟 · 36 字 阅读 →
论文解读

CMI-RewardBench: Evaluating Music Reward Models with Compositional Multimodal Instruction

CMI-RewardBench: Evaluating Music Reward Models with Compositional Multimodal Instruction

 · 更新于 2026-09-09 · 约 1 分钟 · 35 字 阅读 →
论文解读

CoCoEmo: Composable and Controllable Human-Like Emotional TTS via Activation Steering

CoCoEmo: Composable and Controllable Human-Like Emotional TTS via Activation Steering

 · 更新于 2026-09-09 · 约 1 分钟 · 36 字 阅读 →
论文解读

CoLA: Cross-Modal Low-rank Adaptation for Multimodal Downstream Tasks

CoLA: Cross-Modal Low-rank Adaptation for Multimodal Downstream Tasks

 · 更新于 2026-09-09 · 约 1 分钟 · 34 字 阅读 →
论文解读

Convex Low-resource Accent-Robust Language Detection in Speech Recognition

** | 7.5/10

 · 更新于 2026-09-09 · 约 1 分钟 · 78 字 阅读 →
论文解读

DiscoForcing: A Unified Framework for Real-Time Audio-Driven Character Control with Diffusion Forcing

DiscoForcing: A Unified Framework for Real-Time Audio-Driven Character Control with Diffusion Forcing

 · 更新于 2026-09-09 · 约 1 分钟 · 38 字 阅读 →
论文解读

Do Audio LLMs Listen or Read? Analyzing and Mitigating Paralinguistic Failures with VoxParadox

Do Audio LLMs Listen or Read? Analyzing and Mitigating Paralinguistic Failures with VoxParadox

 · 更新于 2026-09-09 · 约 1 分钟 · 39 字 阅读 →
论文解读

DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video Generation

DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video Generation

 · 更新于 2026-09-09 · 约 1 分钟 · 34 字 阅读 →
论文解读

Dual-View Predictive Diffusion: Lightweight Speech Enhancement via Spectrogram-Image Synergy

Dual-View Predictive Diffusion: Lightweight Speech Enhancement via Spectrogram-Image Synergy

 · 更新于 2026-09-09 · 约 1 分钟 · 35 字 阅读 →
论文解读

E-VAds: An E-commerce Short Videos Understanding Benchmark for MLLMs

E-VAds: An E-commerce Short Videos Understanding Benchmark for MLLMs

 · 更新于 2026-09-09 · 约 1 分钟 · 35 字 阅读 →
论文解读

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs

 · 更新于 2026-09-09 · 约 1 分钟 · 34 字 阅读 →
论文解读

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability

 · 更新于 2026-09-09 · 约 1 分钟 · 38 字 阅读 →
论文解读

FakeWorld 1.0: An Omni modal Benchmark for Fake Media and Content

FakeWorld 1.0: An Omni modal Benchmark for Fake Media and Content

 · 更新于 2026-09-09 · 约 1 分钟 · 37 字 阅读 →
论文解读

FoeGlass: When Simple In-Context Learning Is Enough for Red Teaming Audio Deepfake Detectors

FoeGlass: When Simple In-Context Learning Is Enough for Red Teaming Audio Deepfake Detectors

 · 更新于 2026-09-09 · 约 1 分钟 · 39 字 阅读 →
论文解读

From Inpainting to Editing: Unlocking Robust Mask-Free Visual Dubbing via Generative Bootstrapping

From Inpainting to Editing: Unlocking Robust Mask-Free Visual Dubbing via Generative Bootstrapping

 · 更新于 2026-09-09 · 约 1 分钟 · 38 字 阅读 →