论文解读

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment

语音合成 | 9.3/10

 · 更新于 2026-09-30 · 约 13 分钟 · 6263 字 阅读 →
论文解读

Improving acoustic drone detection generalization through pretraining and data augmentation

音频事件检测 | 7.7/10

 · 更新于 2026-09-30 · 约 11 分钟 · 5246 字 阅读 →
论文解读

Latent Space Disentanglement via Activation Steering for Interpretable Attribute Control in Symbolic Music Generation

音乐生成 | 7.3/10

 · 更新于 2026-09-30 · 约 11 分钟 · 5329 字 阅读 →
论文解读

检索增强让高层意图变低层声学:用锚点保检索、用描述做转向的标题投毒

针对检索增强的文本到音乐生成中用检索到的标题改写用户高层提示的依赖,论文提出保留高层锚点并注入低层声学描述的双层标题投毒,使投毒后生成音频与攻击目标类别的 CLAP 相似度从 0.21-0.28 提升到 0.41-0.48,同时与原始用户问题的相似度维持在约 0.30 附近。

 · 更新于 2026-09-30 · 约 21 分钟 · 10208 字 阅读 →
论文解读

MindVoice: Reconstructing Intelligible Speech from Non-invasive Neural Signals with Pretrained Priors

语音合成 | 8.5/10

 · 更新于 2026-09-30 · 约 17 分钟 · 8301 字 阅读 →
论文解读

On the Use of Dereverberation for Acoustic Feedback Cancellation

语音增强 | 6.7/10

 · 更新于 2026-09-30 · 约 8 分钟 · 3988 字 阅读 →
论文解读

OpenSTBench: Beyond Semantic Evaluation for Speech Translation

语音翻译 | 6.0/10

 · 更新于 2026-09-30 · 约 13 分钟 · 6238 字 阅读 →
论文解读

Scaling Conversational Hungarian ASR: The BEA-Dialogue+ Corpus

语音识别 | 7.2/10

 · 更新于 2026-09-30 · 约 12 分钟 · 5792 字 阅读 →
论文解读

Sound effects in media:A comparative analysis of recorded and synthetic samples in live-action and animation

音频生成 | 5.7/10

 · 更新于 2026-09-30 · 约 13 分钟 · 6218 字 阅读 →
论文解读

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue

语音合成 | 8.9/10

 · 更新于 2026-09-30 · 约 17 分钟 · 8091 字 阅读 →
论文解读

Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer

自回归模型 | 6.5/10

 · 更新于 2026-09-30 · 约 17 分钟 · 8023 字 阅读 →
论文解读

UniAudio-Token: Empowering Semantic Speech Tokenizers with General Audio Perception

语音合成 | 10/10

 · 更新于 2026-09-30 · 约 16 分钟 · 7690 字 阅读 →
论文解读

UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion

语音合成 | 8.2/10

 · 更新于 2026-09-30 · 约 16 分钟 · 7971 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-01

共分析 23 篇语音/AI 论文

 · 更新于 2026-09-30 · 约 73 分钟 · 36167 字 阅读 →
论文解读

A Multi-Probe Audit of Clinical-Interview Depression Detection Benchmarks

语音情感识别 | 9.6/10

 · 更新于 2026-09-30 · 约 13 分钟 · 6505 字 阅读 →
论文解读

Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs

语音识别 | 7.2/10

 · 更新于 2026-09-30 · 约 11 分钟 · 5495 字 阅读 →
论文解读

EchoDistill:Alignment Noisy-to-Clean Self-Distillation for Robust Audio LLMs

强化学习 | 9.1/10

 · 更新于 2026-09-30 · 约 13 分钟 · 6317 字 阅读 →
论文解读

MIRAGE: Adaptive Multimodal Gating for Whole-Brain fMRI Encoding

Transformer | 8.2/10

 · 更新于 2026-09-30 · 约 11 分钟 · 5123 字 阅读 →
论文解读

PiAnnotate: A Web Annotation Tool for Piano Fingering, with a Diagnostic Probe

PiAnnotate: A Web Annotation Tool for Piano Fingering, with a Diagnostic Probe

 · 更新于 2026-09-30 · 约 15 分钟 · 7056 字 阅读 →
论文解读

Raon-Speech Technical Report

语音识别 | 6.5/10

 · 更新于 2026-09-30 · 约 11 分钟 · 5464 字 阅读 →