论文解读

Phoneme-Level Deepfake Detection Across Emotional Conditions Using Self-Supervised Embeddings

音频深度伪造检测 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4325 字 阅读 →
论文解读

Toward Structural Multimodal Representations: Specialization, Selection, and Sparsification via Mixture-of-Experts

多模态模型 | 7.0/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6070 字 阅读 →
论文解读

Dimensionality-Aware Anomaly Detection in Learned Representations of Self-Supervised Speech Models

语音识别 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4665 字 阅读 →
论文解读

Virtual Speech Therapist: A Clinician-in-the-Loop AI Speech Therapy Agent for Personalized and Supervised Therapy

语音治疗系统 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4453 字 阅读 →
论文解读

A cross-species neural foundation model for end-to-end speech decoding

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5724 字 阅读 →
论文解读

Alethia: A Foundational Encoder for Voice Deepfakes

音频深度伪造检测 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4189 字 阅读 →
论文解读

AVEX: What Matters for Animal Vocalization Encoding

生物声学 | 7.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5485 字 阅读 →
论文解读

Better Together: Leveraging Unpaired Multimodal Data for Stronger Unimodal Models

音频分类 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4990 字 阅读 →
论文解读

CTC-DRO: Robust Optimization for Reducing Language Disparities in Speech Recognition

语音识别 | 9.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4230 字 阅读 →
论文解读

DiffSDA: Unsupervised Diffusion Sequential Disentanglement Across Modalities

序列解耦 | 8.0/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6401 字 阅读 →
论文解读

Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention

语音分离 | 7.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5777 字 阅读 →
论文解读

FlexiCodec: A Dynamic Neural Audio Codec for Low Frame Rates

语音合成 | 9.0/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5627 字 阅读 →
论文解读

LayerSync: Self-aligning Intermediate Layers

音频生成 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4057 字 阅读 →
论文解读

Learning multimodal dictionary decompositions with group-sparse autoencoders

跨模态检索 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3971 字 阅读 →
论文解读

MAPSS: Manifold-based Assessment of Perceptual Source Separation

模型评估 | 8.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4083 字 阅读 →
论文解读

PACE: Pretrained Audio Continual Learning

音频分类 | 9.0/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5856 字 阅读 →
论文解读

SNAP-UQ: Self-supervised Next-Activation Prediction for Single-Pass Uncertainty in TinyML

音频分类 | 7.5/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6035 字 阅读 →
论文解读

The Deleuzian Representation Hypothesis

模型可解释性 | 8.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4780 字 阅读 →
论文解读

Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio Classification

音频分类 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5053 字 阅读 →
论文解读

AVEX: What Matters for Animal Vocalization Encoding

生物声学 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4577 字 阅读 →