论文解读

The Learning Objective Governs Perceptual Narrowing: A Cross-Lingual, Layer-Wise, Ten-Seed Study of Self-Supervised Speech Encoders

语音属性识别 | 7.0/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6369 字 阅读 →
论文解读

Do Music Foundation Models Embed Pitch in Helical Structure?

音乐理解 | 8.1/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6831 字 阅读 →
论文解读

Integrating Contextual Embeddings into Evaluation of Expressive MIDI Piano Performances

音乐理解 | 7.7/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6828 字 阅读 →
论文解读

Dissecting Sensitivity to Training Language in Self-Supervised Speech Learning Using Neural Audio Codec Tokens

语音识别 | 6.3/10

 · 更新于 2026-09-24 · 约 18 分钟 · 8991 字 阅读 →
论文解读

Less is More: Modality-Decoupling for General AIGC Audio-Video Detection

音频伪造检测 | 7.5/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7402 字 阅读 →
论文解读

AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition

语音情感识别 | 5.6/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5579 字 阅读 →
论文解读

Extracting Voice Styles from Frozen TTS Models via Gradient-Based Inverse Optimization

语音克隆 | 7.9/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5923 字 阅读 →
论文解读

Multi-Phonation Graph Learning with Self-Supervised Speech Embeddings for ALS Detection and Progression Prediction

语音属性识别 | 5.5/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6989 字 阅读 →
论文解读

Self-Supervised Audio Representation Learning for Pediatric Asthma Detection in Emergency Care Using Digital Stethoscope Recordings

音频分类 | 5.0/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5443 字 阅读 →
论文解读

Improving Zero-Shot Phonetic Classification through Language-Agnostic Articulatory Features

语音属性识别 | 5.3/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7242 字 阅读 →
论文解读

Leveraging Gradient Reversal Loss and Multitask Learning for Datasets-Aware Audio Deepfake Detection

语音伪造检测 | 7.3/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6744 字 阅读 →
论文解读

Speech Entrainment in Multi-Party Conversations with a Digital Agent

语音交互 | 5.3/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5599 字 阅读 →
论文解读

Music-JEPA: Learning a World Model of Sound from Action

音乐转录 | 6.2/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6393 字 阅读 →
论文解读

Transforming Keystroke Noise to Text: Self-Supervised Acoustic Eavesdropping Attacks on Keyboards

语音活动检测 | 6.5/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7374 字 阅读 →
论文解读

DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages

语音识别 | 8.1/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7224 字 阅读 →
论文解读

OPOD: On-Policy Omni Distillation

多模态模型 | 7.1/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6936 字 阅读 →
论文解读

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning

音频理解 | 6.8/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7635 字 阅读 →
论文解读

Layer-Wise Decision Fusion for Fake Audio Detection Using XLS-R

音频理解 | 7.5/10

 · 更新于 2026-09-24 · 约 17 分钟 · 8024 字 阅读 →
论文解读

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances

语音编码 | 9.2/10

 · 更新于 2026-09-24 · 约 18 分钟 · 8587 字 阅读 →
论文解读

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection

语音伪造检测 | 6.4/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6832 字 阅读 →