论文解读

Bridging the Age Gap: Towards Detecting Neural Audio Codec Synthesized Elderly Speech Deepfake

语音伪造检测 | 8.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5580 字 阅读 →
论文解读

CoughPhase-CLR: Designing an acoustics-informed foundation model for coughing sound classification

对比学习 | 10/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5062 字 阅读 →
论文解读

ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 23 分钟 · 11361 字 阅读 →
论文解读

How Well Do Self-Supervised Speech Models Encode Age and Gender in Children's Speech? A Layer-Wise Analysis Across Multiple Architectures

自监督学习 | 9/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4728 字 阅读 →
论文解读

MindAlign: Decoding Inner Speech from fMRI Signals via Multimodal Embedding Alignment under Limited Data

语音识别 | 5.8/10

 · 更新于 2026-09-25 · 约 7 分钟 · 3142 字 阅读 →
论文解读

On the Effect of Segmentation Width and Cluster Size on Speech Resynthesis and Continuation in Generative Spoken Language Models

语音合成 | 7.4/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5272 字 阅读 →
论文解读

Online Predictive Coding for Dual-Mode Self-Supervised Speech Model

语音识别 | 7.2/10

 · 更新于 2026-09-25 · 约 21 分钟 · 10439 字 阅读 →
论文解读

Physics-Informed Neural Operator for Speech Production Analysis

语音合成 | 6.7/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4938 字 阅读 →
论文解读

SDP-Codec: A Speaker-Decoupled Speech Codec with Pitch Injection for Low-Bitrate Coding and Zero-Shot Voice Conversion

语音编码 | 7.2/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5887 字 阅读 →
论文解读

Speaker Identity in Non-Verbal Vocalizations: Conditional Distillation and Mixture of Experts Approach

说话人验证 | 9.1/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6090 字 阅读 →
论文解读

Towards Detecting Neural Audio Codec Synthesized Heart Sounds

自监督学习 | 8.7/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6772 字 阅读 →
论文解读

Using Phonological-Level Wav2Vec2 for Mandarin Automatic Mispronunciation Detection and Diagnosis

语音识别 | 8.3/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6199 字 阅读 →
论文解读

A Comparative Study of Pretrained Transformer Models for Quranic ASR: Speech Representations, Label Formats, and Dataset Composition

语音识别 | 7.2/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5716 字 阅读 →
论文解读

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine

语音合成 | 7.4/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5736 字 阅读 →
论文解读

Beyond Speaker Independence: Evaluating Cross-Lingual Acoustic-to-Articulatory Inversion Across Finnish and Russian

自监督学习 | 4.9/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6597 字 阅读 →
论文解读

Cross-Dataset, Age, and Gender Generalization: A Comprehensive Analysis of Fine-Tuning Strategies for Low-Resource Children's ASR

语音识别 | 6.7/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5701 字 阅读 →
论文解读

Improving End-to-End Speech Recognition for Dysarthric Speech through In-Domain Data Augmentation

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 23 分钟 · 11077 字 阅读 →
论文解读

Interpreting Content and Speaker Characteristics in Factorised Self-Supervised Subspaces

语音合成 | 5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4928 字 阅读 →
论文解读

Investigating Human-Model Discrepancies in Speech Quality Assessment via Acoustic and Prosodic Perturbations

语音合成 | 6.9/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7804 字 阅读 →
论文解读

Light-weight Pronunciation Assessment via Discrete Speech Token Surprisal

自监督学习 | 6.4/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6130 字 阅读 →