论文解读

Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS

语音合成 | 10/10

 · 更新于 2026-09-30 · 约 6 分钟 · 2923 字 阅读 →
论文解读

DOA: Training-Free Decoder-Only Attention Policy for Long-Form Simultaneous Translation with SpeechLLMs

语音翻译 | 7.8/10

 · 更新于 2026-09-30 · 约 12 分钟 · 5915 字 阅读 →
论文解读

Escaping the Linearity Trap: Manifold Detours for Black-Box Adversarial Attacks on Singing Audio Deepfake Detection

自监督学习 | 9.7/10

 · 更新于 2026-09-30 · 约 16 分钟 · 7568 字 阅读 →
论文解读

Extracting accent features in spoken Brazilian Portuguese without sociolinguistic labels

口音识别 | 8.3/10

 · 更新于 2026-09-30 · 约 15 分钟 · 7444 字 阅读 →
论文解读

FiPA-SR -- FiLM-Conditioned Perceptually Informed Audio Super-Resolution

生成对抗网络 | 8.1/10

 · 更新于 2026-09-30 · 约 10 分钟 · 4919 字 阅读 →
论文解读

几何主导下如何减出材料:GaMi 的跨模态相减式解耦

针对几何变化压制材料特征的问题,GaMi 用共位毫米波与声学的共享几何一致性做对齐-校准-相减解耦,并以样本间对比抑制残差,在 20 类材料上以 95.2% 的总体准确率显著优于单模态基线,代价是需双模态同步采集与多目标联合训练。

 · 更新于 2026-09-30 · 约 23 分钟 · 11269 字 阅读 →
论文解读

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment

语音合成 | 9.3/10

 · 更新于 2026-09-30 · 约 13 分钟 · 6263 字 阅读 →
论文解读

Improving acoustic drone detection generalization through pretraining and data augmentation

音频事件检测 | 7.7/10

 · 更新于 2026-09-30 · 约 11 分钟 · 5246 字 阅读 →
论文解读

Latent Space Disentanglement via Activation Steering for Interpretable Attribute Control in Symbolic Music Generation

音乐生成 | 7.3/10

 · 更新于 2026-09-30 · 约 11 分钟 · 5329 字 阅读 →
论文解读

检索增强让高层意图变低层声学:用锚点保检索、用描述做转向的标题投毒

针对检索增强的文本到音乐生成中用检索到的标题改写用户高层提示的依赖,论文提出保留高层锚点并注入低层声学描述的双层标题投毒,使投毒后生成音频与攻击目标类别的 CLAP 相似度从 0.21-0.28 提升到 0.41-0.48,同时与原始用户问题的相似度维持在约 0.30 附近。

 · 更新于 2026-09-30 · 约 21 分钟 · 10208 字 阅读 →
论文解读

MindVoice: Reconstructing Intelligible Speech from Non-invasive Neural Signals with Pretrained Priors

语音合成 | 8.5/10

 · 更新于 2026-09-30 · 约 17 分钟 · 8301 字 阅读 →
论文解读

On the Use of Dereverberation for Acoustic Feedback Cancellation

语音增强 | 6.7/10

 · 更新于 2026-09-30 · 约 8 分钟 · 3988 字 阅读 →
论文解读

OpenSTBench: Beyond Semantic Evaluation for Speech Translation

语音翻译 | 6.0/10

 · 更新于 2026-09-30 · 约 13 分钟 · 6238 字 阅读 →
论文解读

Scaling Conversational Hungarian ASR: The BEA-Dialogue+ Corpus

语音识别 | 7.2/10

 · 更新于 2026-09-30 · 约 12 分钟 · 5792 字 阅读 →
论文解读

Sound effects in media:A comparative analysis of recorded and synthetic samples in live-action and animation

音频生成 | 5.7/10

 · 更新于 2026-09-30 · 约 13 分钟 · 6218 字 阅读 →
论文解读

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue

语音合成 | 8.9/10

 · 更新于 2026-09-30 · 约 17 分钟 · 8091 字 阅读 →
论文解读

Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer

自回归模型 | 6.5/10

 · 更新于 2026-09-30 · 约 17 分钟 · 8023 字 阅读 →
论文解读

UniAudio-Token: Empowering Semantic Speech Tokenizers with General Audio Perception

语音合成 | 10/10

 · 更新于 2026-09-30 · 约 16 分钟 · 7690 字 阅读 →
论文解读

UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion

语音合成 | 8.2/10

 · 更新于 2026-09-30 · 约 16 分钟 · 7971 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-01

共分析 23 篇语音/AI 论文

 · 更新于 2026-09-30 · 约 73 分钟 · 36167 字 阅读 →