论文解读

Learning to Align with Unbalanced Optimal Transport in Linguistic Knowledge Transfer for ASR

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4273 字 阅读 →
论文解读

LOTUSDIS: A Thai Far-Field Meeting Corpus for Robust Conversational ASR

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3735 字 阅读 →
论文解读

Low-Resource Speech-Based Early Alzheimers Detection via Cross-Lingual and Few-Shot Transfer Learning

语音生物标志物 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4145 字 阅读 →
论文解读

MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation

语音分离 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4749 字 阅读 →
论文解读

Multi-Layer Attentive Probing Improves Transfer of Audio Representations for Bioacoustics

生物声学 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4488 字 阅读 →
论文解读

Multilingual Supervised Pretraining with Lm-Assisted Decoding for Visual Speech Recognition

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3573 字 阅读 →
论文解读

Perceptual Loss Optimized HRTF Personalization in Spherical Harmonic Domain

空间音频 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5312 字 阅读 →
论文解读

Praxy Voice: Voice-Prompt Recovery + BUPS for Commercial-Class Indic TTS from a Frozen Non-Indic Base at Zero Commercial-Training-Data Cost

语音合成 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4562 字 阅读 →
论文解读

Probing the Hidden Talent of ASR foundation models for L2 English Oral Assessment

预训练 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4262 字 阅读 →
论文解读

Probing Whisper for Dysarthric Speech in Detection and Assessment

语音生物标志物 | 6.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3926 字 阅读 →
论文解读

Quality Assessment of Noisy and Enhanced Speech with Limited Data: UWB-NTIS System for Voicemos 2024

语音质量评估 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5174 字 阅读 →
论文解读

Ranking The Impact of Contextual Specialization in Neural Speech Enhancement

语音增强 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4252 字 阅读 →
论文解读

Representation-Diverse Self-Supervision for Cross-Domain Bioacoustic Learning in Low-Resource Settings

生物声学 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5415 字 阅读 →
论文解读

SAUNA: Song-Level Audio & User-Listening Data Neural Alignment

音乐信息检索 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3885 字 阅读 →
论文解读

SELD-MOHA: A Fine-Tuning Method with the Mixture of Heterogeneous Adapters for Sound Event Localization and Detection

音频事件检测 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5325 字 阅读 →
论文解读

Semantic Anchor Transfer from Short to Long Speech in a Distillation-Based Summarization Framework

语音摘要 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5537 字 阅读 →
论文解读

SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models

音频问答 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5286 字 阅读 →
论文解读

SpeechMapper: Speech-To-Text Embedding Projector for LLMs

语音大模型 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5397 字 阅读 →
论文解读

Stress Prediction from Temporal Emotion Trajectories in Clinical Patient-Physician Conversations

语音情感识别 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4349 字 阅读 →
论文解读

Synthesized Data Selection via Score Distribution Matching for Te Reo Māori Automatic Speech Recognition

语音识别 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4219 字 阅读 →