论文解读

Multilingual Supervised Pretraining with Lm-Assisted Decoding for Visual Speech Recognition

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3573 字 阅读 →
论文解读

Neuromamba: Adaptive Frequency Filtering with a Pyramid Mamba for sEEG-driven Speech Synthesis

语音合成 | 8.0/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5962 字 阅读 →
论文解读

One Model–Three Tasks: Discovering a Shared Winning Ticket for Low-Complexity Audio Intelligence

音频分类 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3730 字 阅读 →
论文解读

Polynomial Mixing for Efficient Self-Supervised Speech Encoders

语音识别 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4653 字 阅读 →
论文解读

Praxy Voice: Voice-Prompt Recovery + BUPS for Commercial-Class Indic TTS from a Frozen Non-Indic Base at Zero Commercial-Training-Data Cost

语音合成 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4562 字 阅读 →
论文解读

Prompt-Guided Mixture-of-Experts for Robust Multimodal Sentiment Analysis with Missing Modalities

语音情感识别 | 8.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4723 字 阅读 →
论文解读

Quantifying Speaker Embedding Phonological Rule Interactions in Accented Speech Synthesis

语音合成 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4134 字 阅读 →
论文解读

Ranking The Impact of Contextual Specialization in Neural Speech Enhancement

语音增强 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4252 字 阅读 →
论文解读

Representation-Diverse Self-Supervision for Cross-Domain Bioacoustic Learning in Low-Resource Settings

生物声学 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5415 字 阅读 →
论文解读

Scaling Ambiguity: Augmenting Human Annotation in Speech Emotion Recognition with Audio-Language Models

语音情感识别 | 6.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4289 字 阅读 →
论文解读

Self-Supervised Note Tracking and Multi-Pitch Estimation Via Reconstruction-Based Learning

多音高估计 音符跟踪 | 8.5/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9289 字 阅读 →
论文解读

Sequence-Level Unsupervised Training in Speech Recognition: A Theoretical Study

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 7 分钟 · 3356 字 阅读 →
论文解读

SSVD-O: Parameter-Efficient Fine-Tuning with Structured SVD for Speech Recognition

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5250 字 阅读 →
论文解读

Synthesized Data Selection via Score Distribution Matching for Te Reo Māori Automatic Speech Recognition

语音识别 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4219 字 阅读 →
论文解读

TAGARELA - A Portuguese Speech Dataset from Podcasts

语音识别 语音合成 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3967 字 阅读 →
论文解读

Taming Audio VAEs via Target-KL Regularization

音频生成 | 6.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5919 字 阅读 →
论文解读

Task Vector in TTS: Toward Emotionally Expressive Dialectal Speech Synthesis

语音合成 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4678 字 阅读 →
论文解读

Three Seconds is Sufficient: A Multi-Pronged Framework for Model-Based Speaker Adaptation in ASR Under Data-Scarce Conditions

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4318 字 阅读 →
论文解读

TICL: Text-Embedding KNN for Speech in-Context Learning Unlocks Speech Recognition Abilities of Large Multimodal Models

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4710 字 阅读 →
论文解读

TMD-TTS: A Unified Tibetan Multi-Dialect Text-to-Speech Framework for Ü-Tsang, Amdo and Kham Speech Dataset Generation

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4601 字 阅读 →