论文解读

BFCL Audio: An Audio Function Calling Evaluation for Large Language Models

语音交互 | 7.7/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7260 字 阅读 →
论文解读

Bioacoustic Geolocation: Species Sounds as Geographic Signals

音频理解 | 5.8/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6365 字 阅读 →
论文解读

Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models

语音合成 | 8/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7969 字 阅读 →
论文解读

Characterizing the Predictive Impact of Modalities with Supervised Latent-Variable Modeling

Characterizing the Predictive Impact of Modalities with Supervised Latent-Variable Modeling

 · 更新于 2026-09-24 · 约 13 分钟 · 6181 字 阅读 →
论文解读

CMI-RewardBench: Evaluating Music Reward Models with Compositional Multimodal Instruction

音乐生成 | 6.4/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4833 字 阅读 →
论文解读

CoCoEmo: Composable and Controllable Human-Like Emotional TTS via Activation Steering

语音合成 | 7.1/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6544 字 阅读 →
论文解读

CoLA: Cross-Modal Low-rank Adaptation for Multimodal Downstream Tasks

音视频理解 | 8.3/10

 · 更新于 2026-09-24 · 约 17 分钟 · 8159 字 阅读 →
论文解读

ConsMSA: Semantic Distribution Consistency Learning for Multimodal Sentiment Analysis

多模态模型 | 6.1/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5246 字 阅读 →
论文解读

Convex Low-resource Accent-Robust Language Detection in Speech Recognition

语音识别 | 6/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7467 字 阅读 →
论文解读

Decoupling The "What" and "Where" With Polar Coordinate Positional Embedding

音乐生成 | 7.8/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6360 字 阅读 →
论文解读

DiscoForcing: A Unified Framework for Real-Time Audio-Driven Character Control with Diffusion Forcing

音乐生成 | 8/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6820 字 阅读 →
论文解读

Do Audio LLMs Listen or Read? Analyzing and Mitigating Paralinguistic Failures with VoxParadox

语音属性识别 | 8/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6346 字 阅读 →
论文解读

DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video Generation

音视频生成 | 8/10

 · 更新于 2026-09-24 · 约 20 分钟 · 9827 字 阅读 →
论文解读

Dual-View Predictive Diffusion: Lightweight Speech Enhancement via Spectrogram-Image Synergy

语音增强 | 8.4/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5900 字 阅读 →
论文解读

E-VAds: An E-commerce Short Videos Understanding Benchmark for MLLMs

音视频问答 | 6.9/10

 · 更新于 2026-09-24 · 约 17 分钟 · 8438 字 阅读 →
论文解读

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs

音视频理解 | 6.3/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6514 字 阅读 →
论文解读

Efficient Distributed MLLM Training with Cornstarch

音视频理解 | 7/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5041 字 阅读 →
论文解读

Efficient Multi-modal Dataset Distillation via Analytic Parameter Matching

对比学习 | 7.2/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6217 字 阅读 →
论文解读

Efficient, Property-Aligned Fan-Out Retrieval via RL-Compiled Diffusion

音乐检索 | 4.7/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7918 字 阅读 →
论文解读

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability

语音合成 | 6.6/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6527 字 阅读 →