论文解读

CRAFT: Critic-Refined Adaptive Key-Frame Targeting for Multimodal Video Question Answering

多模态问答 | 6.6/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8953 字 阅读 →
论文解读

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action

语音对话系统 | 7.8/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9025 字 阅读 →
论文解读

Music of Changing Lines: Toward a Culturally Situated Approach to the I-Ching

音乐生成 | 5.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6040 字 阅读 →
论文解读

PlanRAG-Audio: Planning and Retrieval Augmented Generation for Long-form Audio Understanding

长音频理解 | 7.4/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7334 字 阅读 →
论文解读

Stage-adaptive Token Selection for Efficient Omni-modal LLMs

多模态模型 | 7.7/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8274 字 阅读 →
论文解读

Can Large Language Models Reliably Correct Errors in Low-Resource ASR? A Contamination-Aware Case Study on West Frisian

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9491 字 阅读 →
论文解读

OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding

视频理解 | 7.3/10

 · 更新于 2026-09-25 · 约 18 分钟 · 9001 字 阅读 →
论文解读

Towards Trust Calibration in Socially Interactive Agents: Investigating Gendered Multimodal Behaviors Generation with LLMs

社交智能体 | 5.9/10

 · 更新于 2026-09-25 · 约 21 分钟 · 10069 字 阅读 →
论文解读

Can Large Language Models Imitate Human Speech for Clinical Assessment? LLM-Driven Data Augmentation for Cognitive Score Prediction

语音生物标志物 | 6/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8390 字 阅读 →
论文解读

Can We Trust AI-Inferred User States. A Psychometric Framework for Validating the Reliability of Users States Classification by LLMs in Operational Environments

模型评估 | 6/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7392 字 阅读 →
论文解读

Improving Automatic Speech Recognition for Speakers Treated for Oral Cancer using Data Augmentation and LLM Error Correction

语音识别 | 6/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8703 字 阅读 →
论文解读

FutureSim: Replaying World Events to Evaluate Adaptive Agents

基准测试 | 7.6/10

 · 更新于 2026-09-25 · 约 20 分钟 · 9806 字 阅读 →
论文解读

MediaClaw: Multimodal Intelligent-Agent Platform Technical Report

多模态模型 | 3.3/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7378 字 阅读 →
论文解读

Refining Pseudo-Audio Prompts with Speech-Text Alignment for Text-Only Domain Adaptation in LLM-Based ASR

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8162 字 阅读 →
论文解读

PresentAgent-2: Towards Generalist Multimodal Presentation Agents

生成模型 | 6.5/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7708 字 阅读 →
论文解读

Text2Score: Generating Sheet Music From Textual Prompts

乐谱生成 | 7.0/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9318 字 阅读 →
论文解读

Mind the Pause: Disfluency-Aware Objective Tuning for Multilingual Speech Correction with LLMs

语音编辑 | 6.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5409 字 阅读 →
论文解读

Speech-based Psychological Crisis Assessment using LLMs

语音情感识别 | 5.8/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8451 字 阅读 →
论文解读

MIST: Multimodal Interactive Speech-based Tool-calling Conversational Assistants for Smart Homes

MIST: Multimodal Interactive Speech-based Tool-calling Conversational Assistants for Smart Homes

 · 更新于 2026-09-25 · 约 14 分钟 · 7004 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-11

共分析 12 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 41 分钟 · 20158 字 阅读 →