论文解读

Light-weight Pronunciation Assessment via Discrete Speech Token Surprisal

自监督学习 | 6.4/10

 · 更新于 2026-09-07 · 约 13 分钟 · 6130 字 阅读 →
论文解读

Low-Burden Data Augmentation for Dysarthric ASR via Zero-Shot Voice Cloning

语音识别 | 8.7/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5408 字 阅读 →
论文解读

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model

语音合成 | 5.7/10

 · 更新于 2026-09-07 · 约 8 分钟 · 3893 字 阅读 →
论文解读

MixProLAP: Mixture-Induced Uncertainty Modeling for Probabilistic Language-Audio Pretraining

音频检索 | 5.7/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5203 字 阅读 →
论文解读

NEST: Narrative Event Structures in Time for Long Video Understanding

NEST: Narrative Event Structures in Time for Long Video Understanding

 · 更新于 2026-09-07 · 约 10 分钟 · 4950 字 阅读 →
论文解读

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models

多模态模型 | 8.1/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5119 字 阅读 →
论文解读

Personalized Keyword Spotting for User-Defined Keywords Leveraging Text-Independent Speaker Verification

说话人验证 | 8.6/10

 · 更新于 2026-09-07 · 约 16 分钟 · 7723 字 阅读 →
论文解读

PhysDrift: Bridging the Embodiment Gap in Humanoid Co-Speech Motion Generation

语音合成 | 7.4/10

 · 更新于 2026-09-07 · 约 14 分钟 · 6816 字 阅读 →
论文解读

Pitch Spelling Jazz Lead Sheets, Solo Transcriptions, Classical Piano and Monophonic Scores

Pitch Spelling Jazz Lead Sheets, Solo Transcriptions, Classical Piano and Monophonic Scores

 · 更新于 2026-09-07 · 约 15 分钟 · 7476 字 阅读 →
论文解读

PolSeT: Polish Semantics of Timbre Dataset

PolSeT: Polish Semantics of Timbre Dataset

 · 更新于 2026-09-07 · 约 7 分钟 · 3358 字 阅读 →
论文解读

PrefSQA: Pairwise Preference Prediction for Speech Quality Assessment and the Critical Role of High Quality Datasets

语音质量评估 | 7.3/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5828 字 阅读 →
论文解读

Prismriver: Formalization of Music Theory and Algorithmic Composition in Lean 4

Prismriver: Formalization of Music Theory and Algorithmic Composition in Lean 4

 · 更新于 2026-09-07 · 约 9 分钟 · 4353 字 阅读 →
论文解读

ReNikud: Audio-Supervised Hebrew Grapheme-to-Phoneme Conversion

语音合成 | 6.2/10

 · 更新于 2026-09-07 · 约 16 分钟 · 7884 字 阅读 →
论文解读

Repurposing a Speech Classifier for Guided Diffusion-Based Speech Generation

语音合成 | 7.9/10

 · 更新于 2026-09-07 · 约 14 分钟 · 6654 字 阅读 →
论文解读

RIVET: Robust Idempotent Voice Attribute Editing

语音转换 | 8/10

 · 更新于 2026-09-07 · 约 9 分钟 · 4451 字 阅读 →
论文解读

S-JEPA : Soft Clustering Anchors for Self-Supervised Speech Representation Learning

语音识别 | 8.7/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5785 字 阅读 →
论文解读

Segment-Level Mandarin Chinese Speech-Based Cognitive Impairment Detection via an Autoencoder with Contrastive Learning

对比学习 | 6.5/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5574 字 阅读 →
论文解读

Stuttering Classification and Segmentation with Attention-Based Multiple Instance Learning

Stuttering Classification and Segmentation with Attention-Based Multiple Instance Learning

 · 更新于 2026-09-07 · 约 11 分钟 · 5207 字 阅读 →
论文解读

Systematic Study of Dysarthric Speech Recognition: Spectral Features and Acoustic Models

语音识别 | 8.3/10

 · 更新于 2026-09-07 · 约 10 分钟 · 4610 字 阅读 →
论文解读

Time-Unconditional Generative Speech Enhancement via Autonomous Rectified Flow

语音增强 | 7.0/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5796 字 阅读 →