论文解读

G-MaP-SE: Guided Speech Enhancement via GMM-Based Prior Matching

语音增强 | 9.3/10

 · 更新于 2026-09-09 · 约 11 分钟 · 5176 字 阅读 →
论文解读

HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis

语音合成 | 9/10

 · 更新于 2026-09-09 · 约 14 分钟 · 6999 字 阅读 →
论文解读

Is Text All You Need? Text as a Universal Information Bottleneck for Speech LLMs

语音识别 | 7.6/10

 · 更新于 2026-09-09 · 约 11 分钟 · 5311 字 阅读 →
论文解读

Liberating LLM Capabilities in Full-Duplex Speech Models

多模态模型 | 8.7/10

 · 更新于 2026-09-09 · 约 17 分钟 · 8443 字 阅读 →
论文解读

MeanVC 2: Robust Low-Latency Streaming Zero-Shot Voice Conversion

语音转换 | 6.9/10

 · 更新于 2026-09-09 · 约 23 分钟 · 11480 字 阅读 →
论文解读

MeCo: One-Step MeanFlow-based Corrector for Multi-Channel Speech Separation

语音分离 | 8.4/10

 · 更新于 2026-09-09 · 约 12 分钟 · 5695 字 阅读 →
论文解读

Multi-View Speech Representation Learning for Parkinson's Disease Detection Using Context-guided Cross-modal Attention

自监督学习 | 7.9/10

 · 更新于 2026-09-09 · 约 12 分钟 · 6004 字 阅读 →
论文解读

NüshuVoice: Reviving the Voice of Endangered Nüshu with Pitch-Aware Text-to-Speech

语音合成 | 7/10

 · 更新于 2026-09-09 · 约 11 分钟 · 5298 字 阅读 →
论文解读

OmniMem: Perturbation-aware Memory Compression for Streaming Audio-Visual LLMs

高效推理 | 8/10

 · 更新于 2026-09-09 · 约 11 分钟 · 5135 字 阅读 →
论文解读

On Low-Bit Quantization Errors in Speaker Verification: Diagnostic and Mitigation

说话人验证 | 6.6/10

 · 更新于 2026-09-09 · 约 11 分钟 · 5313 字 阅读 →
论文解读

OpenBibleTTS: Large-Scale Speech Resources and TTS Models for Low-Resource Languages

语音合成 | 8/10

 · 更新于 2026-09-09 · 约 14 分钟 · 6802 字 阅读 →
论文解读

Overcoming Decoder Inconsistencies in Whisper for Dravidian and Low-Resource Languages

语音识别 | 6.2/10

 · 更新于 2026-09-09 · 约 13 分钟 · 6234 字 阅读 →
论文解读

Paediatric-HGNN: A Hybrid Heterogeneous Graph Neural Network for Detecting Disfluency in Children's Speech via Multiscale Acoustic Fusion

语音合成 | 6.5/10

 · 更新于 2026-09-09 · 约 15 分钟 · 7186 字 阅读 →
论文解读

Parameter-Efficient Continual Learning for Automatic Speech Recognition

语音识别 | 8.1/10

 · 更新于 2026-09-09 · 约 11 分钟 · 5336 字 阅读 →
论文解读

Predictive Fixed-Filter Active Noise Control (PFANC) Using Convolutional Recurrent Neural Networks for Dynamic Noises

Predictive Fixed-Filter Active Noise Control (PFANC) Using Convolutional Recurrent Neural Networks for Dynamic Noises

 · 更新于 2026-09-09 · 约 12 分钟 · 5760 字 阅读 →
论文解读

Probing Token Spaces under Generator Shift in AI-Generated Music Detection

音频编码 | 9/10

 · 更新于 2026-09-09 · 约 12 分钟 · 6007 字 阅读 →
论文解读

Quality-Diversity Search in Sound Generation: Investigating Innovation Engines for Audio Exploration

Quality-Diversity Search in Sound Generation: Investigating Innovation Engines for Audio Exploration

 · 更新于 2026-09-09 · 约 24 分钟 · 12020 字 阅读 →
论文解读

Rethinking Depth: A study of the Recursive-Transformer for Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-09 · 约 12 分钟 · 5625 字 阅读 →
论文解读

SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation

音频生成 | 7/10

 · 更新于 2026-09-09 · 约 12 分钟 · 5816 字 阅读 →
论文解读

Sound Field Interpolation Using Physics-Informed Extreme Learning Machine with Pre-Training

语音增强 | 5.3/10

 · 更新于 2026-09-09 · 约 12 分钟 · 5857 字 阅读 →