论文解读

Amplifying Membership Signal Through Chained Regeneration

生成模型 | 6.6/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5952 字 阅读 →
论文解读

ASR-Agnostic Multimodal Spectrotemporal Modeling for Early Dementia Detection

多模态模型 | 7.4/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8321 字 阅读 →
论文解读

Attacking UTMOS: Probing the Robustness of a Speech Quality Assessment Model

语音质量评估 | 8.6/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5401 字 阅读 →
论文解读

AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation

语音合成 | 6.5/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6668 字 阅读 →
论文解读

BEST-RQ-2: Contextualize-Then-Predict, a Two-Step Approach for Self-Supervised Audio Representations

语音识别 | 6.9/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4866 字 阅读 →
论文解读

Beyond Binary Instrument QA: Probing Instrument Grounding in Music Audio-Language Models

音频分类 | 7.6/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4548 字 阅读 →
论文解读

Beyond Cross-Reconstruction: Probing-Based Disentanglement Evaluation for Acoustic Teleportation Codecs

语音编码 | 8.1/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5500 字 阅读 →
论文解读

Building a Multimodal Dataset of Academic Paper for Keyword Extraction

语音识别 | 5.2/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5542 字 阅读 →
论文解读

Building an ASR Solution for Training and Assessing Children's Reading

语音识别 | 8.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5557 字 阅读 →
论文解读

Detecting Audio Deepfakes on the Edge:Lightweight SSL-Based Detection in a Browser Plugin

Detecting Audio Deepfakes on the Edge:Lightweight SSL-Based Detection in a Browser Plugin

 · 更新于 2026-09-06 · 约 13 分钟 · 6339 字 阅读 →
论文解读

Dilemmadata: On the Interoperability of Heterogeneous Roman Numeral Datasets

数据集 | 10/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4937 字 阅读 →
论文解读

Enhancing BEST-RQ Pseudo-Label Quality through Online Refinement for Automatic Speech Recognition

语音识别 | 8.6/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5117 字 阅读 →
论文解读

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model

语音合成 | 7.2/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6790 字 阅读 →
论文解读

Gated Multi-Graph Fusion via Graph Attention Networks for Alzheimer's Disease Detection

语音情感识别 | 5.2/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5054 字 阅读 →
论文解读

How Bilingual Are SSL Speech Models? Cross-Lingual Probing of Articulatory Encoding with Finnish and Russian EMA

自监督学习 | 5.8/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6340 字 阅读 →
论文解读

Improving multichannel speech enhancement through accurate room-acoustic simulations

语音增强 | 6.8/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5624 字 阅读 →
论文解读

Is Natural Always Appropriate? Investigating Naturalness and Appropriateness Across Different Domains for TTS Evaluation

语音合成 | 7.2/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7931 字 阅读 →
论文解读

Linguistic Bias Mitigation for Spoofing Detection via Gradient Reversal and A Variational Information Bottleneck

自监督学习 | 8.6/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5540 字 阅读 →
论文解读

Listening Between the Lines: Joint Learning of ASR Embeddings and LLM-Augmented Linguistics for Dementia Detection

数据增强 | 7/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6468 字 阅读 →
论文解读

LOPA: Enhancing Spoken Language Assessment via Latent Ordinal Prototype Alignment

低资源 | 6.2/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6248 字 阅读 →