论文解读

Polynomial Mixing for Efficient Self-Supervised Speech Encoders

语音识别 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4653 字 阅读 →
论文解读

Praxy Voice: Voice-Prompt Recovery + BUPS for Commercial-Class Indic TTS from a Frozen Non-Indic Base at Zero Commercial-Training-Data Cost

语音合成 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4562 字 阅读 →
论文解读

Prompt-Guided Mixture-of-Experts for Robust Multimodal Sentiment Analysis with Missing Modalities

语音情感识别 | 8.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4723 字 阅读 →
论文解读

Quantifying Speaker Embedding Phonological Rule Interactions in Accented Speech Synthesis

语音合成 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4134 字 阅读 →
论文解读

Ranking The Impact of Contextual Specialization in Neural Speech Enhancement

语音增强 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4252 字 阅读 →
论文解读

Representation-Diverse Self-Supervision for Cross-Domain Bioacoustic Learning in Low-Resource Settings

生物声学 | 7.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5415 字 阅读 →
论文解读

Scaling Ambiguity: Augmenting Human Annotation in Speech Emotion Recognition with Audio-Language Models

语音情感识别 | 6.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4289 字 阅读 →
论文解读

Self-Supervised Note Tracking and Multi-Pitch Estimation Via Reconstruction-Based Learning

多音高估计 音符跟踪 | 8.5/10

 · 更新于 2026-09-06 · 约 19 分钟 · 9289 字 阅读 →
论文解读

Sequence-Level Unsupervised Training in Speech Recognition: A Theoretical Study

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 7 分钟 · 3356 字 阅读 →
论文解读

SSVD-O: Parameter-Efficient Fine-Tuning with Structured SVD for Speech Recognition

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5250 字 阅读 →
论文解读

Synthesized Data Selection via Score Distribution Matching for Te Reo Māori Automatic Speech Recognition

语音识别 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4219 字 阅读 →
论文解读

TAGARELA - A Portuguese Speech Dataset from Podcasts

语音识别 语音合成 | 7.0/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3967 字 阅读 →
论文解读

Taming Audio VAEs via Target-KL Regularization

音频生成 | 6.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5919 字 阅读 →
论文解读

Task Vector in TTS: Toward Emotionally Expressive Dialectal Speech Synthesis

语音合成 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4678 字 阅读 →
论文解读

Three Seconds is Sufficient: A Multi-Pronged Framework for Model-Based Speaker Adaptation in ASR Under Data-Scarce Conditions

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4318 字 阅读 →
论文解读

TICL: Text-Embedding KNN for Speech in-Context Learning Unlocks Speech Recognition Abilities of Large Multimodal Models

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4710 字 阅读 →
论文解读

TMD-TTS: A Unified Tibetan Multi-Dialect Text-to-Speech Framework for Ü-Tsang, Amdo and Kham Speech Dataset Generation

语音合成 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4601 字 阅读 →
论文解读

Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6217 字 阅读 →
论文解读

Towards Lightweight Adaptation of Speech Enhancement Models in Real-World Environments

语音增强 | 8.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4520 字 阅读 →
论文解读

Towards Orthographically-Informed Evaluation of Speech Recognition Systems for Indian Languages

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3966 字 阅读 →