论文解读

Prosody-Guided Harmonic Attention for Phase-Coherent Neural Vocoding in the Complex Spectrum

语音合成 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4866 字 阅读 →
论文解读

PSTalker: Realistic 3D Talking Head Synthesis via a Semantic-Aware Audio-Driven Point-Based Shape

说话人合成 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4433 字 阅读 →
论文解读

Real-Time Streaming MEL Vocoding with Generative Flow Matching

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4120 字 阅读 →
论文解读

Robust Online Overdetermined Independent Vector Analysis Based on Bilinear Decomposition

语音分离 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4197 字 阅读 →
论文解读

SFM-TTS: Lightweight and Rapid Speech Synthesis with Flexible Shortcut Flow Matching

语音合成 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5280 字 阅读 →
论文解读

Shortcut Flow Matching for Speech Enhancement: Step-Invariant Flows via Single Stage Training

语音增强 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4631 字 阅读 →
论文解读

Spring Reverb Emulation with Hybrid Gated Convolutional Networks and State Space Models

音频生成 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5227 字 阅读 →
论文解读

Stereophonic Acoustic Echo Cancellation Using an Improved Affine Projection Algorithm with Adaptive Multiple Sub-Filters

语音增强 | 6.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4348 字 阅读 →
论文解读

Str-DiffSep: Streamable Diffusion Model for Speech Separation

语音分离 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4807 字 阅读 →
论文解读

Stream-Voice-Anon: Enhancing Utility of Real-Time Speaker Anonymization Via Neural Audio Codec and Language Models

语音匿名化 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5467 字 阅读 →
论文解读

Synchronous Secondary Path Modeling and Kronecker-Factorized Adaptive Algorithm for Multichannel Active Noise Control

主动噪声控制 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4919 字 阅读 →
论文解读

T-Cache: Fast Inference For Masked Generative Transformer-Based TTS Via Prompt-Aware Feature Caching

语音合成 | 9.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4509 字 阅读 →
论文解读

T-Mimi: A Transformer-Based Mimi Decoder for Real-Time On-Phone TTS

语音合成 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3841 字 阅读 →
论文解读

Time-Domain Synthesis of Virtual Sound Source Within Personalized Sound Zone using a Linear Loudspeaker Array

空间音频 | 8.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3765 字 阅读 →
论文解读

Towards Real-Time Generative Speech Restoration with Flow-Matching

语音增强 | 6.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4390 字 阅读 →
论文解读

UJCodec: An End-to-end Unet-Style Codec for Joint Speech Compression and Enhancement

语音增强 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4672 字 阅读 →
论文解读

VChangeCodec: An Ultra Low-Complexity Neural Speech Codec with Built-In Voice Changer for Customized Real-Time Communication

语音转换 语音增强 | 8.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5237 字 阅读 →
论文解读

WhisperPipe: A Resource-Efficient Streaming Architecture for Real-Time Automatic Speech Recognition

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 7 分钟 · 3139 字 阅读 →
论文解读

Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation

音视频 | 8.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5626 字 阅读 →
论文解读

Opening the Design Space: Two Years of Performance with Intelligent Musical Instruments

音乐生成 | 6.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4910 字 阅读 →