论文解读

SPARCLE: SPeaker-aware Aligned Representations via Contrastive Language Embeddings

语音合成 | 5.8/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5960 字 阅读 →
论文解读

TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue

语音交互 | 7/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5205 字 阅读 →
论文解读

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization

音视频理解 | 8.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5320 字 阅读 →
论文解读

Disentangling Speaker and Language Effects in Cross-Lingual Speaker Verification for Iberian Languages

说话人验证 | 5.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5518 字 阅读 →
论文解读

BEST-RQ-2: Contextualize-Then-Predict, a Two-Step Approach for Self-Supervised Audio Representations

语音识别 | 6.9/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4866 字 阅读 →
论文解读

Beyond Binary Instrument QA: Probing Instrument Grounding in Music Audio-Language Models

音频分类 | 7.6/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4548 字 阅读 →
论文解读

Enhancing BEST-RQ Pseudo-Label Quality through Online Refinement for Automatic Speech Recognition

语音识别 | 8.6/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5117 字 阅读 →
论文解读

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model

语音合成 | 7.2/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6790 字 阅读 →
论文解读

How Bilingual Are SSL Speech Models? Cross-Lingual Probing of Articulatory Encoding with Finnish and Russian EMA

自监督学习 | 5.8/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6340 字 阅读 →
论文解读

Linguistic Bias Mitigation for Spoofing Detection via Gradient Reversal and A Variational Information Bottleneck

自监督学习 | 8.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5540 字 阅读 →
论文解读

Listening Between the Lines: Joint Learning of ASR Embeddings and LLM-Augmented Linguistics for Dementia Detection

数据增强 | 7/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6468 字 阅读 →
论文解读

LuxEmo: Expressive Text-to-Speech Corpus for Luxembourgish

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6067 字 阅读 →
论文解读

Probing-Guided Layer Selection from Self-Supervised Speech Models for Generalizable Audio Deepfake Detection

集成学习 | 7.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6363 字 阅读 →
论文解读

What Counts as an Error? Dual-Reference Benchmarking for Atypical ASR

语音识别 | 7.3/10

 · 更新于 2026-09-25 · 约 25 分钟 · 12422 字 阅读 →
论文解读

AMR: Adaptive Modality Routing for Multimodal Polyglot Speaker Identification

说话人识别 | 7.8/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6161 字 阅读 →
论文解读

Child-Centric Voice Anonymization in Single and Multi-Speaker Speech via Domain-Adapted SSL Models

语音匿名化 | 7.2/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6716 字 阅读 →
论文解读

Clustering Unsupervised Representations as Defense against Poisoning Attacks on Speech Commands Classification System

自监督学习 | 6.5/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7576 字 阅读 →
论文解读

OLIVE: View-Augmented Latent Prediction with Waveform Reconstruction for Speech SSL

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7900 字 阅读 →
论文解读

SIGMA: Saliency-Guided Sparse Mask Attacks for Speech Emotion Recognition

语音情感识别 | 7.1/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5240 字 阅读 →
论文解读

VeRe-Flow: Guiding Flow Matching toward Clean Speech via Velocity Contrastive Regularization and Representation Alignment for Noise-Robust Bandwidth Expansion

语音增强 | 7.7/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4813 字 阅读 →