论文解读

BiRQ: Bi-Level Self-Labeling Random Quantization for Self-Supervised Speech Recognition

语音识别 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4898 字 阅读 →
论文解读

Break-the-Beat! Controllable MIDI-to-Drum audio synthesis

音乐生成 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4419 字 阅读 →
论文解读

Bridging the Semantic Gap: Cross-Attentive Fusion for Joint Acoustic-Semantic Speech Quality Assessment

语音质量评估 | 8.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5819 字 阅读 →
论文解读

CASTELLA: Long Audio Dataset with Captions and Temporal Boundaries

音频检索 | 8.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3821 字 阅读 →
论文解读

Combining SSL Speech Features, Contextual Transformers and Mamba Models for Realistic Audio Spoofing Detection

音频深度伪造检测 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4304 字 阅读 →
论文解读

Contrastive Timbre Representations for Musical Instrument And Synthesizer Retrieval

音频检索 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3791 字 阅读 →
论文解读

Cross-Lingual Interleaving for Speech Language Models

语音大模型 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5122 字 阅读 →
论文解读

DisContSE: Single-Step Diffusion Speech Enhancement based on Joint Discrete and Continuous Embeddings

语音增强 | 8.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5379 字 阅读 →
论文解读

Do Foundational Audio Encoders Understand Music Structure?

音乐信息检索 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3755 字 阅读 →
论文解读

Does the Pre-Training of an Embedding Influence its Encoding of Age?

语音生物标志物 | 7.0/10

 · 更新于 2026-09-25 · 约 7 分钟 · 3128 字 阅读 →
论文解读

Domain Partitioning Meets Parameter-Efficient Fine-Tuning: A Novel Method for Improved Language-Queried Audio Source Separation

音频分离 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5493 字 阅读 →
论文解读

Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems

语音对话系统 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4864 字 阅读 →
论文解读

Efficient Audio-Visual Inference Via Token Clustering And Modality Fusion

音频问答 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4866 字 阅读 →
论文解读

Efficient Depression Detection from Speech via Language-Independent Prompt-Driven Reprogramming

语音生物标志物 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4393 字 阅读 →
论文解读

Emotional Dimension Control in Language Model-Based Text-To-Speech: Spanning a Broad Spectrum of Human Emotions

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4379 字 阅读 →
论文解读

Enhancing Speaker Verification with w2v-BERT 2.0 and Knowledge Distillation Guided Structured Pruning

说话人验证 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4903 字 阅读 →
论文解读

Enhancing Speech Intelligibility Prediction for Hearing Aids with Complementary Speech Foundation Model Representations

语音增强 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3521 字 阅读 →
论文解读

Exploring How Audio Effects Alter Emotion with Foundation Models

音乐理解 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4321 字 阅读 →
论文解读

FUSEMOS: Perceptual Evaluation of Text-to-Music Generation with Dual-Encoder Fusion and Ranking-Aware Composite Loss

音乐生成 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4701 字 阅读 →
论文解读

Gen-SER: When the Generative Model Meets Speech Emotion Recognition

语音情感识别 | 6.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3718 字 阅读 →