论文解读

Escaping the Linearity Trap: Manifold Detours for Black-Box Adversarial Attacks on Singing Audio Deepfake Detection

自监督学习 | 9.7/10

 · 更新于 2026-09-09 · 约 16 分钟 · 7568 字 阅读 →
论文解读

Extracting accent features in spoken Brazilian Portuguese without sociolinguistic labels

口音识别 | 8.3/10

 · 更新于 2026-09-09 · 约 15 分钟 · 7444 字 阅读 →
论文解读

FiPA-SR -- FiLM-Conditioned Perceptually Informed Audio Super-Resolution

生成对抗网络 | 8.1/10

 · 更新于 2026-09-09 · 约 10 分钟 · 4919 字 阅读 →
论文解读

GaMi: Geometry-Agnostic Material Identification via Cross-Modal Subtractive Disentanglement

GaMi: Geometry-Agnostic Material Identification via Cross-Modal Subtractive Disentanglement

 · 更新于 2026-09-09 · 约 14 分钟 · 6859 字 阅读 →
论文解读

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment

语音合成 | 9.3/10

 · 更新于 2026-09-09 · 约 13 分钟 · 6263 字 阅读 →
论文解读

Improving acoustic drone detection generalization through pretraining and data augmentation

音频事件检测 | 7.7/10

 · 更新于 2026-09-09 · 约 11 分钟 · 5246 字 阅读 →
论文解读

Latent Space Disentanglement via Activation Steering for Interpretable Attribute Control in Symbolic Music Generation

音乐生成 | 7.3/10

 · 更新于 2026-09-09 · 约 11 分钟 · 5329 字 阅读 →
论文解读

Mental Damage: Caption Poisoning Attacks on Retrieval-Augmented Text-to-Music Generation

音乐生成 | 5.6/10

 · 更新于 2026-09-09 · 约 12 分钟 · 5959 字 阅读 →
论文解读

MindVoice: Reconstructing Intelligible Speech from Non-invasive Neural Signals with Pretrained Priors

语音合成 | 8.5/10

 · 更新于 2026-09-09 · 约 17 分钟 · 8301 字 阅读 →
论文解读

On the Use of Dereverberation for Acoustic Feedback Cancellation

语音增强 | 6.7/10

 · 更新于 2026-09-09 · 约 8 分钟 · 3988 字 阅读 →
论文解读

OpenSTBench: Beyond Semantic Evaluation for Speech Translation

语音翻译 | 6.0/10

 · 更新于 2026-09-09 · 约 13 分钟 · 6238 字 阅读 →
论文解读

Scaling Conversational Hungarian ASR: The BEA-Dialogue+ Corpus

语音识别 | 7.2/10

 · 更新于 2026-09-09 · 约 12 分钟 · 5792 字 阅读 →
论文解读

Sound effects in media:A comparative analysis of recorded and synthetic samples in live-action and animation

音频生成 | 5.7/10

 · 更新于 2026-09-09 · 约 13 分钟 · 6218 字 阅读 →
论文解读

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue

语音合成 | 8.9/10

 · 更新于 2026-09-09 · 约 17 分钟 · 8091 字 阅读 →
论文解读

Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer

自回归模型 | 6.5/10

 · 更新于 2026-09-09 · 约 17 分钟 · 8023 字 阅读 →
论文解读

UniAudio-Token: Empowering Semantic Speech Tokenizers with General Audio Perception

语音合成 | 10/10

 · 更新于 2026-09-09 · 约 16 分钟 · 7690 字 阅读 →
论文解读

UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion

语音合成 | 8.2/10

 · 更新于 2026-09-09 · 约 16 分钟 · 7971 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-01

共分析 23 篇语音/AI 论文

 · 更新于 2026-09-09 · 约 73 分钟 · 36167 字 阅读 →
论文解读

A Multi-Probe Audit of Clinical-Interview Depression Detection Benchmarks

语音情感识别 | 9.6/10

 · 更新于 2026-09-09 · 约 13 分钟 · 6505 字 阅读 →
论文解读

Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs

语音识别 | 7.2/10

 · 更新于 2026-09-09 · 约 11 分钟 · 5495 字 阅读 →