论文解读

PC-MCL: Patient-Consistent Multi-Cycle Learning with Multi-Label Bias Correction for Respiratory Sound Classification

音频分类 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 5005 字 阅读 →
论文解读

Peeking Into the Future for Contextual Biasing

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5086 字 阅读 →
论文解读

Phonological Tokenizer: Prosody-Aware Phonetic Token Via Multi-Objective Fine-Tuning with Differentiable K-Means

语音表示学习 | 8.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5452 字 阅读 →
论文解读

Probing Whisper for Dysarthric Speech in Detection and Assessment

语音生物标志物 | 6.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3926 字 阅读 →
论文解读

Proficiency-Aware Adaptation and Data Augmentation for Robust L2 ASR

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3992 字 阅读 →
论文解读

PROST-LLM: Progressively Enhancing the Speech-to-Speech Translation Capability in LLMs

语音翻译 | 7.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5541 字 阅读 →
论文解读

Purification Before Fusion: Toward Mask-Free Speech Enhancement for Robust Audio-Visual Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4265 字 阅读 →
论文解读

Reference-Aware SFM Layers for Intrusive Intelligibility Prediction

语音评估 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4990 字 阅读 →
论文解读

SEP-ST: Incorporating Speech Entity Prompt Into Large Language Models for Speech Translation

语音翻译 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3975 字 阅读 →
论文解读

Session-Level Spoken Language Assessment with A Multimodal Foundation Model Via Multi-Target Learning

语音评估 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4959 字 阅读 →
论文解读

Shared Representation Learning for Reference-Guided Targeted Sound Detection

音频事件检测 | 8.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4764 字 阅读 →
论文解读

Stress Prediction from Temporal Emotion Trajectories in Clinical Patient-Physician Conversations

语音情感识别 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4349 字 阅读 →
论文解读

Task-Oriented Sound Privacy Preservation for Sound Event Detection Via End-to-End Adversarial Multi-Task Learning

音频事件检测 | 7.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5739 字 阅读 →
论文解读

Text2Move: Text-To-Moving Sound Generation via Trajectory Prediction and Temporal Alignment

空间音频 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4225 字 阅读 →
论文解读

Tokenchain: A Discrete Speech Chain via Semantic Token Modeling

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6054 字 阅读 →
论文解读

Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6217 字 阅读 →
论文解读

Transfer Learning for Paediatric Sleep Apnoea Detection using Physiology-Guided Acoustic Models

音频分类 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4463 字 阅读 →
论文解读

Triad: Tri-Head with Auxiliary Duplicating Permutation Invariant Training for Multi-Task Sound Event Localization and Detection

音频事件检测 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4020 字 阅读 →
论文解读

TTA: Transcribe, Translate and Alignment for Cross-Lingual Speech Representation

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5261 字 阅读 →
论文解读

Vioptt: Violin Technique-Aware Transcription from Synthetic Data Augmentation

音乐信息检索 | 6.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4985 字 阅读 →