论文解读

Mitigating Stethoscope-Induced Shortcuts in Respiratory Sound Classification under Federated Domain Generalization with Causality-Inspired Interventions

音频分类 | 8.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5537 字 阅读 →
论文解读

MusTBENCH: Benchmarking and Advancing Temporal Grounding in Music LLMs

音乐生成 | 7.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5633 字 阅读 →
论文解读

Native Audio-Visual Alignment for Generation

音频生成 | 7.8/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6977 字 阅读 →
论文解读

OmniInteract: Benchmarking Real-World Streaming Interaction for Real-Time Omnimodal Assistants

语音识别 | 7.8/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6454 字 阅读 →
论文解读

State-Anchored Complete-View Distillation for Robust Conversational Multimodal Emotion Recognition

语音情感识别 | 8/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7970 字 阅读 →
论文解读

VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents

语音合成 | 7.9/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7179 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-29

共分析 20 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 57 分钟 · 28306 字 阅读 →
论文解读

A Conflict-Aware Penalty and Statistical Loss Framework for Balancing Modalities and Enhancing Stability in Multimodal Sentiment Analysis

多模态模型 | 6.8/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6974 字 阅读 →
论文解读

AgenticVBench: Can AI Agents Complete Real-World Post-Production Tasks?

基准测试 | 7.0/10

 · 更新于 2026-09-06 · 约 19 分钟 · 9207 字 阅读 →
论文解读

Dasheng AudioGen: A Unified Model for Generating Coherent Audio Scenes from Text

音频生成 | 8.6/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7601 字 阅读 →
论文解读

Do Audio LLMs Listen or Read? Analyzing and Mitigating Paralinguistic Failures with VoxParadox

语音情感识别 | 6.8/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5436 字 阅读 →
论文解读

EigeNet: Geometry-Informed Multi-Modal Learning for Few-shot Novel View RIR Prediction

多模态模型 | 8/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5702 字 阅读 →
论文解读

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini

语音识别 | 7.9/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5387 字 阅读 →
论文解读

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation

语音生成 | 9.9/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6224 字 阅读 →
论文解读

SMILE-Next: Teaching Large Language Models to Detect, Classify, and Reason about Laughter

语音情感识别 | 8.7/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5028 字 阅读 →
论文解读

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts

语音合成 | 8/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4782 字 阅读 →
论文解读

Why We Need Speech to Evaluate Speech Translation

语音翻译 | 8.3/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6500 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-28

共分析 30 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 84 分钟 · 42014 字 阅读 →
论文解读

A Multimodal Framework for Dementia Detection via Linguistic and Acoustic Representation Learning

多模态模型 | 7.7/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6784 字 阅读 →
论文解读

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models

多模态模型 | 9.7/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6285 字 阅读 →