论文解读

谐音不是错字:让普通话 ASR 在纠错与保留笑点之间踩刹车

语音识别 | 5.7/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4240 字 阅读 →
论文解读

Multi-Modal Generative Fuzzy System: Fuzzy Inference Guided Large Model Interactive Question Answering Framework

音视频问答 | 5.9/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5449 字 阅读 →
论文解读

MazzikaAI: A knowledge-based performance-to-prompt compiler for real-time Arabic maqam accompaniment with a streaming text-to-music model

音乐生成 | 6.6/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7239 字 阅读 →
论文解读

The MADRS Pipeline: Supporting Depression Assessment in Clinical Trials

医疗音频 | 5.8/10

 · 更新于 2026-09-24 · 约 18 分钟 · 8722 字 阅读 →
论文解读

Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis

音频交互 | 7.0/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6613 字 阅读 →
论文解读

Harness TTS: Towards Context-Aware Expressive Speech Synthesis with Harness Layer

语音合成 | 6.2/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6532 字 阅读 →
论文解读

When to Use Extra Context: Evidence-Grounded Terminology Adaptation for Simultaneous Speech Translation

语音翻译 | 6.7/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6074 字 阅读 →
论文解读

Anamnesis: An Open-Source Platform for Large-Scale Backstory-Conditioned Survey Simulation

提示学习 | 6.1/10

 · 更新于 2026-09-24 · 约 17 分钟 · 8074 字 阅读 →
论文解读

CHARM: Charge Calibration and Acoustic Rescue for LLM-based Multimodal Sarcasm Detection

提示学习 | 8.8/10

 · 更新于 2026-09-24 · 约 21 分钟 · 10446 字 阅读 →
论文解读

ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions

语音合成 | 7.6/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6226 字 阅读 →
论文解读

FoeGlass: Simple In-Context Learning Is Enough for Red Teaming Audio Deepfake Detectors

语音伪造检测 | 6.8/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6644 字 阅读 →
论文解读

Zero-Shot Rankability: Revealing Latent Ordinal Structure in Multimodal Large Language Models via Language

音视频理解 | 6.8/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6730 字 阅读 →
论文解读

ZEBRA: Zero-Shot Entropy-Regularized Prompt Learning for Base-to-Novel Generalization in Audio-Language Models

音频分类 | 7.1/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6499 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-07-01

共分析 35 篇语音/AI 论文

 · 更新于 2026-09-24 · 约 102 分钟 · 50963 字 阅读 →
论文解读

Acoustic Prompting via Stage-wise Modulation for Few-Shot Learning in Audio Language Models

音频分类 | 8.3/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5504 字 阅读 →
论文解读

Decoding while Adapting: Zero-Shot Online Speaker Adaptation via Audio-Textual Prompts for Elderly Speech Recognition

语音识别 | 6.8/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6816 字 阅读 →
论文解读

Ontology Memory-Augmented ASR Correction for Long Text-Speech Interleaved Conversations

语音识别 | 9.6/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6360 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-12

共分析 27 篇语音/AI 论文

 · 更新于 2026-09-24 · 约 77 分钟 · 38140 字 阅读 →
论文解读

Acoustic Cue Alignment in Audio Language Models for Speech Emotion Recognition

语音情感识别 | 7.8/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5285 字 阅读 →
论文解读

MixFake: Benchmarking and Enhancing Audio Deepfake Detection in Diverse Real-world Mixed Audio

音频深度伪造检测 | 10/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7102 字 阅读 →