论文解读

FC-TTS: Style and Timbre Control in Zero-Shot Text-to-Speech with Disentangled Speech Representations

语音合成 | 8.4/10

 · 更新于 2026-10-01 · 约 13 分钟 · 6292 字 阅读 →
论文解读

Hidden in Plain Tokens: Simply Robust, Gradient-Free Watermark for Synthetic Audio

音频水印 | 6.2/10

 · 更新于 2026-10-01 · 约 13 分钟 · 6460 字 阅读 →
论文解读

Multilingual Phonological Feature Recognition with Self-Supervised Speech Models

语音识别 | 7.7/10

 · 更新于 2026-10-01 · 约 12 分钟 · 5589 字 阅读 →
论文解读

Music Transcription with (Almost) No Supervision

音乐转录 | 10/10

 · 更新于 2026-10-01 · 约 12 分钟 · 5872 字 阅读 →
论文解读

Proactive for Uncertainty: Cause-Aware Error Diagnosis and Interactive Clarification for Spoken Dialogue Systems

语音识别 | 9.6/10

 · 更新于 2026-10-01 · 约 17 分钟 · 8124 字 阅读 →
论文解读

Rethinking Continual Learning for Speech and Audio: A Representation-Centric Taxonomy and Open Problems

语音识别 | 6.0/10

 · 更新于 2026-10-01 · 约 9 分钟 · 4440 字 阅读 →
论文解读

Rubato: Transcribing Piano Music with Timestamps

音乐转录 | 7.5/10

 · 更新于 2026-10-01 · 约 13 分钟 · 6346 字 阅读 →
论文解读

Score-Agnostic Structure Analysis in Large-Scale Performance Datasets

音乐信息检索 | 4.1/10

 · 更新于 2026-10-01 · 约 11 分钟 · 5407 字 阅读 →
论文解读

SpongeBob: Sync-Aware Harmonious Audio-Visual Generative Editing

语音编辑 | 8.6/10

 · 更新于 2026-10-01 · 约 13 分钟 · 6408 字 阅读 →
论文解读

StrTransformer: Source-Wise Structured Transformers for Unsupervised Blind Source Recovery

StrTransformer: Source-Wise Structured Transformers for Unsupervised Blind Source Recovery

 · 更新于 2026-10-01 · 约 11 分钟 · 5030 字 阅读 →
论文解读

Subspace Track-before-Detect for Passive Multi-Target Tracking with Unknown Emitted Signals

声源定位 | 5.5/10

 · 更新于 2026-10-01 · 约 11 分钟 · 5308 字 阅读 →
论文解读

Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation

Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation

 · 更新于 2026-10-01 · 约 12 分钟 · 5656 字 阅读 →
论文解读

Thaka at KSAA-2026 Task 2: Regularized Fine-Tuning for Arabic Speech Diacritization

语音增强 | 5.5/10

 · 更新于 2026-10-01 · 约 12 分钟 · 5587 字 阅读 →
论文解读

The Symmetric Location Problem: a Song of Efficiency and Robustness

The Symmetric Location Problem: a Song of Efficiency and Robustness

 · 更新于 2026-10-01 · 约 15 分钟 · 7297 字 阅读 →
论文解读

Time Segmented Beamforming via Dynamic Programming: Theory and Implementation

实时处理 | 7.7/10

 · 更新于 2026-10-01 · 约 14 分钟 · 6521 字 阅读 →
论文解读

Toward Native Multimodal Modeling: A Roadmap

多模态模型 | 10/10

 · 更新于 2026-10-01 · 约 10 分钟 · 4994 字 阅读 →
论文解读

Toward Natural Emotional Text-To-Speech System with Fine-Grained Non-Verbal Expression Control

语音合成 | 5.0/10

 · 更新于 2026-10-01 · 约 11 分钟 · 5339 字 阅读 →
论文解读

Ultra-Low-Bitrate Mel-Spectrogram-based Neural Speech Coding with Flow-Matching-based Refinement and Vocoding-driven Reconstruction

语音编码 | 9.9/10

 · 更新于 2026-10-01 · 约 17 分钟 · 8226 字 阅读 →
论文解读

WaveNeXt 2: ConvNeXt-Based Fast Neural Vocoders With Residual Denoising and Sub-Modeling for GAN and Diffusion Models

语音合成 | 9.4/10

 · 更新于 2026-10-01 · 约 11 分钟 · 5426 字 阅读 →
论文解读

Zero-Shot Parkinson's Disease Detection from Speech: Comparing Large Audio and Language Models

大语言模型 | 5/10

 · 更新于 2026-10-01 · 约 14 分钟 · 6830 字 阅读 →