论文解读

Better Together: Leveraging Unpaired Multimodal Data for Stronger Unimodal Models

多模态模型 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4906 字 阅读 →
论文解读

LayerSync: Self-aligning Intermediate Layers

生成模型 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4835 字 阅读 →
论文解读

MAPSS: Manifold-based Assessment of Perceptual Source Separation

语音分离 | 8.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4535 字 阅读 →
论文解读

MindMix: A Multimodal Foundation Model for Auditory Perception Decoding via Deep Neural-Acoustic Alignment

音频检索 | 8.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5061 字 阅读 →
论文解读

The Deleuzian Representation Hypothesis

模型评估 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4035 字 阅读 →
论文解读

TVTSyn: Content-Synchronous Time-Varying Timbre for Streaming Voice Conversion and Anonymization

语音转换 | 7.5/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6178 字 阅读 →
论文解读

Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio Classification

音频分类 | 9.0/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5834 字 阅读 →
论文解读

Advancing automatic speech recognition using feature fusion with self-supervised learning features: A case study on Fearless Steps Apollo corpus

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5309 字 阅读 →
论文解读

Beyond Acoustic Sparsity and Linguistic Bias: A Prompt-Free Paradigm for Mispronunciation Detection and Diagnosis

发音错误检测 | 8.5/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6451 字 阅读 →
论文解读

Identifying and typifying demographic unfairness in phoneme-level embeddings of self-supervised speech recognition models

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4560 字 阅读 →
论文解读

PSP: An Interpretable Per-Dimension Accent Benchmark for Indic Text-to-Speech

语音合成 | 9.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4333 字 阅读 →
论文解读

Similarity Choice and Negative Scaling in Supervised Contrastive Learning for Deepfake Audio Detection

音频深度伪造检测 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4571 字 阅读 →
论文解读

SPG-Codec: Exploring the Role and Boundaries of Semantic Priors in Ultra-Low-Bitrate Neural Speech Coding

语音合成 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4115 字 阅读 →
论文解读

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation

语音情感识别 模型评估 | 7.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5334 字 阅读 →
论文解读

A Parameter-Efficient Multi-Scale Convolutional Adapter for Synthetic Speech Detection

A Parameter-Efficient Multi-Scale Convolutional Adapter for Synthetic Speech Detection

 · 更新于 2026-09-06 · 约 9 分钟 · 4431 字 阅读 →
论文解读

A Study of Data Selection Strategies for Pre-Training Self-Supervised Speech Models

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4254 字 阅读 →
论文解读

A Superb-Style Benchmark of Self-Supervised Speech Models for Audio Deepfake Detection

音频深度伪造检测 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4149 字 阅读 →
论文解读

A Task-Aware Dual-Level Self-Supervised Learning Method for Effective Sound Event Detection

音频事件检测 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4113 字 阅读 →
论文解读

Advanced modeling of interlanguage speech intelligibility benefit with L1-L2 multi-task learning using differentiable K-means for accent-robust discrete token-based ASR

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4144 字 阅读 →
论文解读

Advancing Semi-Supervised Child Speech Recognition with Omni-Temporal Classification under Label Noise

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4912 字 阅读 →