每日研究速递

语音/音乐/音频论文速递 2026-05-15

共分析 20 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 83 分钟 · 41092 字 阅读 →
论文解读

EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents

语音对话系统 | 8.0/10

 · 更新于 2026-09-25 · 约 20 分钟 · 9925 字 阅读 →
论文解读

GeoBuildBench: A Benchmark for Interactive and Executable Geometry Construction from Natural Language

几何推理 | 7.0/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7034 字 阅读 →
论文解读

PresentAgent-2: Towards Generalist Multimodal Presentation Agents

生成模型 | 6.5/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7708 字 阅读 →
论文解读

Senses Wide Shut: A Representation-Action Gap in Omnimodal LLMs

模型评估 | 8.0/10

 · 更新于 2026-09-25 · 约 22 分钟 · 10630 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-14

共分析 16 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 62 分钟 · 30769 字 阅读 →
论文解读

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation

多模态模型评估 | 5.5/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9061 字 阅读 →
论文解读

MMTB: Evaluating Terminal Agents on Multimedia-File Tasks

基准测试 | 60/10

 · 更新于 2026-09-25 · 约 20 分钟 · 9917 字 阅读 →
论文解读

The Deepfakes We Missed: We Built Detectors for a Threat That Didn't Arrive

深度伪造检测 | 6.5/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7431 字 阅读 →
论文解读

Too Good to Be True: A Study on Modern Automatic Speech Recognition for the Evaluation of Speech Enhancement

语音增强 | 6.6/10

 · 更新于 2026-09-25 · 约 21 分钟 · 10496 字 阅读 →
论文解读

Towards Fine-Grained Multi-Dimensional Speech Understanding: Data Pipeline, Benchmark, and Model

Towards Fine-Grained Multi-Dimensional Speech Understanding: Data Pipeline, Benchmark, and Model

 · 更新于 2026-09-25 · 约 18 分钟 · 8956 字 阅读 →
论文解读

Evaluating the Expressive Appropriateness of Speech in Rich Contexts

语音质量评估 | 7.2/10

 · 更新于 2026-09-25 · 约 20 分钟 · 9625 字 阅读 →
论文解读

FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries

音频检索 | 6.0/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9517 字 阅读 →
论文解读

Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search

基准测试 | 6.0/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7890 字 阅读 →
论文解读

Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization

基准测试 | 6.5/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8876 字 阅读 →
论文解读

RADAR Challenge 2026: Robust Audio Deepfake Recognition under Media Transformations

音频深度伪造检测 | 6.0/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7522 字 阅读 →
论文解读

Responsible Benchmarking of Fairness for Automatic Speech Recognition

语音识别 | 5.0/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7970 字 阅读 →
论文解读

Do Joint Audio-Video Generation Models Understand Physics?

Do Joint Audio-Video Generation Models Understand Physics?

 · 更新于 2026-09-25 · 约 18 分钟 · 8531 字 阅读 →
论文解读

MIST: Multimodal Interactive Speech-based Tool-calling Conversational Assistants for Smart Homes

MIST: Multimodal Interactive Speech-based Tool-calling Conversational Assistants for Smart Homes

 · 更新于 2026-09-25 · 约 14 分钟 · 7004 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-11

共分析 12 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 41 分钟 · 20158 字 阅读 →