论文解读

Frame-Aligned Fusion of Canary and WavLM for Non-Intrusive Intelligibility Prediction of Hearing-Aid-Processed Speech

语音质量评估 | 10/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6946 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-25

共分析 19 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 52 分钟 · 25789 字 阅读 →
论文解读

A Survey of Audio Reasoning in Multimodal Foundation Models

音频推理 | 7.7/10

 · 更新于 2026-09-06 · 约 18 分钟 · 8530 字 阅读 →
论文解读

CRAFT: Critic-Refined Adaptive Key-Frame Targeting for Multimodal Video Question Answering

多模态问答 | 6.6/10

 · 更新于 2026-09-06 · 约 18 分钟 · 8953 字 阅读 →
论文解读

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action

语音对话系统 | 7.8/10

 · 更新于 2026-09-06 · 约 19 分钟 · 9025 字 阅读 →
论文解读

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation

基准测试 | 8.1/10

 · 更新于 2026-09-06 · 约 20 分钟 · 9524 字 阅读 →
论文解读

Stage-adaptive Token Selection for Efficient Omni-modal LLMs

多模态模型 | 7.7/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8274 字 阅读 →
论文解读

CounterFlow: A Two-Phase Inference-Time Sampling for Counterfactual Video Foley Generation

音频生成 | 6/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7631 字 阅读 →
论文解读

EMO-BOOST: Emotion-Augmented Audio-Visual Features for Improved Generalization in Deepfake Detection

音频深度伪造检测 | 7.2/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8115 字 阅读 →
论文解读

GroupAffect-4: A Multimodal Dataset of Four-Person Collaborative Interaction

数据集 | 6.8/10

 · 更新于 2026-09-06 · 约 19 分钟 · 9307 字 阅读 →
论文解读

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation

基准测试 | 6.5/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8197 字 阅读 →
论文解读

OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding

视频理解 | 7.3/10

 · 更新于 2026-09-06 · 约 18 分钟 · 9001 字 阅读 →
论文解读

When Vision Speaks for Sound

音视频 | 7.7/10

 · 更新于 2026-09-06 · 约 19 分钟 · 9514 字 阅读 →
论文解读

Acoustic Interference: A New Paradigm Weaponizing Acoustic Latent Semantic for Universal Jailbreak against Large Audio Language Models

音频安全 | 8.7/10

 · 更新于 2026-09-06 · 约 19 分钟 · 9033 字 阅读 →
论文解读

Audio-Image Cross-Modal Retrieval with Onomatopoeic Images

音频检索 | 7/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6984 字 阅读 →
论文解读

CodeBind: Decoupled Representation Learning for Multimodal Alignment with Unified Compositional Codebook

多模态模型 | 8.6/10

 · 更新于 2026-09-06 · 约 21 分钟 · 10221 字 阅读 →
论文解读

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation

音视频 | 7.3/10

 · 更新于 2026-09-06 · 约 19 分钟 · 9061 字 阅读 →
论文解读

SIREM: Speech-Informed MRI Reconstruction with Learned Sampling

医学图像重建 | 7.3/10

 · 更新于 2026-09-06 · 约 18 分钟 · 8537 字 阅读 →
论文解读

VISAFF: Speaker-Centered Visual Affective Feature Learning for Emotion Recognition in Conversation

对话情感识别 | 7.4/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6752 字 阅读 →
论文解读

Sound Sparks Motion: Audio and Text Tuning for Video Editing

视频编辑 | 5.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5597 字 阅读 →