论文解读

Beyond Generative Decoding: Discriminative Hidden-State Readout from a Native Omni-Modal LLM for Multimodal Sentiment Analysis

多模态模型 | 5.3/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6789 字 阅读 →
论文解读

Beyond Text Following: Repairable Arbitration Reversals in Audio-Language Models

音频问答 | 6.4/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6582 字 阅读 →
论文解读

Forgive or forget: Understanding the context of hate in audio retrieval systems

音频检索 | 7.4/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6360 字 阅读 →
论文解读

M2S-AVSR: Modality-aware Multi-view Self-supervised Representation for Robust Audio-Visual Speech Recognition

语音识别 | 9/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6425 字 阅读 →
论文解读

MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models

语音识别 | 8.4/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6696 字 阅读 →
论文解读

To Be Multimodal or Not to Be: Query-Adaptive Audio-Visual Person Retrieval via Active Modality Detection

说话人识别 | 6.8/10

 · 更新于 2026-09-25 · 约 27 分钟 · 13289 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-05

共分析 47 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 127 分钟 · 63374 字 阅读 →
论文解读

DetectZoo: A Unified Toolkit for AI-Generated Content Detection Across Text, Audio, and Image Modalities

多模态模型 | 9.3/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6131 字 阅读 →
论文解读

Entity Binding Failures in Speech LLM Reasoning: Diagnosis and Chain-of-Thought Intervention

语音问答 | 7.5/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8342 字 阅读 →
论文解读

Multilingual Long-Form Speech Instruction Following: KIT's Submission to IWSLT 2026

语音识别 | 10/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6280 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-04

共分析 22 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 60 分钟 · 29896 字 阅读 →
论文解读

A Pocket Offline Model for Simultaneous Speech Translation as CUNI Submission to IWSLT 2026

语音翻译 | 6.8/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6364 字 阅读 →
论文解读

Before Fusion, Ask What to Keep: Contextual Calibration of Multimodal Signals

语音情感识别 | 7.2/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6279 字 阅读 →
论文解读

Benchmarking Speech-to-Speech Translation Models

语音合成 | 8.7/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5395 字 阅读 →
论文解读

Cosmos 3: Omnimodal World Models for Physical AI

音频生成 | 10/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8607 字 阅读 →
论文解读

Foley-Omni: A Unified Multimodal Generation Model from Task-Level Audio Synthesis to Complete Video Soundtrack Generation

音频生成 | 7/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6522 字 阅读 →
论文解读

MoDAl: Self-Supervised Neural Modality Discovery via Decorrelation for Speech Neuroprosthesis

自监督学习 | 6.6/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5302 字 阅读 →
论文解读

OmniHalluc-L: Counterfactual Benchmarking and Modality-Perturbation Reliability Calibration for Long-Form Omni Hallucination

多模态模型 | 7.8/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6119 字 阅读 →
论文解读

SegTune: Structured and Fine-Grained Control for Song Generation

音乐生成 | 8.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6788 字 阅读 →
论文解读

SketchSong: Hierarchical Song Generation with Sketch Planning and Fine-Grained Multi-Track Modeling

音乐生成 | 8.6/10

 · 更新于 2026-09-25 · 约 26 分钟 · 12623 字 阅读 →