论文解读

Synchronization and Turn-Taking in Full-Duplex Speech Dialogue Models

语音对话系统 | 7.6/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7848 字 阅读 →
论文解读

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation

基准测试 | 6.5/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8197 字 阅读 →
论文解读

CodeBind: Decoupled Representation Learning for Multimodal Alignment with Unified Compositional Codebook

多模态模型 | 8.6/10

 · 更新于 2026-09-25 · 约 21 分钟 · 10221 字 阅读 →
论文解读

ARIA: A Diagnostic Framework for Music Training Data Attribution

音乐生成 | 6.1/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9163 字 阅读 →
论文解读

Can We Trust AI-Inferred User States. A Psychometric Framework for Validating the Reliability of Users States Classification by LLMs in Operational Environments

模型评估 | 6/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7392 字 阅读 →
论文解读

ViMU: Benchmarking Video Metaphorical Understanding

基准测试 | 8.1/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7731 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-17

共分析 2 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 9 分钟 · 4469 字 阅读 →
论文解读

A Benchmark for Early-stage Parkinson's Disease Detection from Speech

语音生物标志物 | 7.2/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7457 字 阅读 →
论文解读

From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents

模型评估 | 6.3/10

 · 更新于 2026-09-25 · 约 22 分钟 · 10971 字 阅读 →
论文解读

Senses Wide Shut: A Representation-Action Gap in Omnimodal LLMs

模型评估 | 8.0/10

 · 更新于 2026-09-25 · 约 22 分钟 · 10630 字 阅读 →
论文解读

jina-embeddings-v5-omni: Text-Geometry-Preserving Multimodal Embeddings via Frozen-Tower Composition

多模态检索 | 7.5/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8909 字 阅读 →
论文解读

The SMC Blind Spot: A Failure Mode Analysis of State-of-the-Art Beat Tracking

节拍跟踪 | 7.4/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6909 字 阅读 →
论文解读

Too Good to Be True: A Study on Modern Automatic Speech Recognition for the Evaluation of Speech Enhancement

语音增强 | 6.6/10

 · 更新于 2026-09-25 · 约 21 分钟 · 10496 字 阅读 →
论文解读

Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search

基准测试 | 6.0/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7890 字 阅读 →
论文解读

Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization

基准测试 | 6.5/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8876 字 阅读 →
论文解读

Responsible Benchmarking of Fairness for Automatic Speech Recognition

语音识别 | 5.0/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7970 字 阅读 →
论文解读

Towards Trustworthy Audio Deepfake Detection: A Systematic Framework for Diagnosing and Mitigating Gender Bias

音频深度伪造检测 | 6.5/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8536 字 阅读 →
论文解读

Do Joint Audio-Video Generation Models Understand Physics?

Do Joint Audio-Video Generation Models Understand Physics?

 · 更新于 2026-09-25 · 约 18 分钟 · 8531 字 阅读 →
论文解读

Evaluating voice anonymisation using similarity rank disclosure

Evaluating voice anonymisation using similarity rank disclosure

 · 更新于 2026-09-25 · 约 15 分钟 · 7465 字 阅读 →
论文解读

Sparse Autoencoders as Plug-and-Play Firewalls for Adversarial Attack Detection in VLMs

Sparse Autoencoders as Plug-and-Play Firewalls for Adversarial Attack Detection in VLMs

 · 更新于 2026-09-25 · 约 18 分钟 · 8609 字 阅读 →