论文解读

Advanced modeling of interlanguage speech intelligibility benefit with L1-L2 multi-task learning using differentiable K-means for accent-robust discrete token-based ASR

语音识别 | 7.0/10

 · 更新于 2026-09-11 · 约 9 分钟 · 4144 字 阅读 →
论文解读

Advancing LLM-Based Multi-Channel Multi-Speaker Speech Recognition with Global Cross-Channel Attention and Sentence-Ordered First-In First-Out Serialized Output Training

语音识别 | 7.5/10

 · 更新于 2026-09-11 · 约 11 分钟 · 5058 字 阅读 →
论文解读

Advancing Semi-Supervised Child Speech Recognition with Omni-Temporal Classification under Label Noise

语音识别 | 7.5/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4912 字 阅读 →
论文解读

Advancing Speech Summarization in Multi-Modal LLMs with Reinforcement Learning

音频问答 | 7.0/10

 · 更新于 2026-09-11 · 约 9 分钟 · 4272 字 阅读 →
论文解读

Advancing Speech Understanding in Speech-Aware Language Models with GRPO

语音问答 | 7.0/10

 · 更新于 2026-09-11 · 约 7 分钟 · 3396 字 阅读 →
论文解读

Adversarial Defense via Generative Speech Enhancement Module

语音增强 对抗防御 | 7.5/10

 · 更新于 2026-09-11 · 约 8 分钟 · 3619 字 阅读 →
论文解读

Adversarial Fine-Tuning on Speech Foundation Model with Vulnerable Attention Consistency Regularization for Robust Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-11 · 约 9 分钟 · 4509 字 阅读 →
论文解读

Adversarial Rivalry Learning for Music Classification

音乐分类 | 6.5/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4912 字 阅读 →
论文解读

Affect-Jigsaw: Integrating Core and Peripheral Emotions for Harmonious Fine-Grained Multimodal Emotion Recognition

语音情感识别 | 8.0/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4583 字 阅读 →
论文解读

AFT: An Exemplar-Free Class Incremental Learning Method for Environmental Sound Classification

音频分类 | 7.0/10

 · 更新于 2026-09-11 · 约 9 分钟 · 4408 字 阅读 →
论文解读

AI-Generated Music Detection in Broadcast Monitoring

音频深度伪造检测 | 7.0/10

 · 更新于 2026-09-11 · 约 9 分钟 · 4218 字 阅读 →
论文解读

Ailive Mixer: A Deep Learning Based Zero Latency Automatic Music Mixer for Live Music Performances

音乐混合 | 7.0/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4873 字 阅读 →
论文解读

AISHELL6-Whisper: A Chinese Mandarin Audio-Visual Whisper Speech Dataset with Speech Recognition Baselines

语音识别 | 8.3/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4735 字 阅读 →
论文解读

Aligning Generative Speech Enhancement with Perceptual Feedback

语音增强 | 7.5/10

 · 更新于 2026-09-11 · 约 12 分钟 · 5533 字 阅读 →
论文解读

Aligning Language Models for Lyric-to-Melody Generation with Rule-Based Musical Constraints

音乐生成 | 7.5/10

 · 更新于 2026-09-11 · 约 9 分钟 · 4356 字 阅读 →
论文解读

ALMA-Chor: Leveraging Audio-Lyric Alignment with Mamba for Chorus Detection

音乐信息检索 | 7.0/10

 · 更新于 2026-09-11 · 约 9 分钟 · 4379 字 阅读 →
论文解读

AMBER2: Dual Ambiguity-Aware Emotion Recognition Applied to Speech and Text

语音情感识别 | 8.0/10

 · 更新于 2026-09-11 · 约 9 分钟 · 4014 字 阅读 →
论文解读

AmbiDrop: Array-Agnostic Speech Enhancement Using Ambisonics Encoding and Dropout-Based Learning

语音增强 | 7.0/10

 · 更新于 2026-09-11 · 约 62 分钟 · 30885 字 阅读 →
论文解读

AMBISONIC-DML: A Benchmark Dataset for Dynamic Higher-Order Ambisonics Music with Motion-Aligned Stems

数据集 | 7.5/10

 · 更新于 2026-09-11 · 约 9 分钟 · 4335 字 阅读 →
论文解读

An Anomaly-Aware and Audio-Enhanced Dual-Pathway Framework for Alzheimer’s Disease Progression Classification

语音生物标志物 | 7.0/10

 · 更新于 2026-09-11 · 约 12 分钟 · 5640 字 阅读 →