论文解读

Improving Binaural Distance Estimation in Reverberant Rooms Through Contrastive And Multi-Task Learning

声源定位 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4718 字 阅读 →
论文解读

In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word level timestamp predictions

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4573 字 阅读 →
论文解读

InstructAudio: Unified Speech and Music Generation with Natural Language Instruction

语音合成 | 7.5/10

 · 更新于 2026-09-06 · 约 19 分钟 · 9232 字 阅读 →
论文解读

It Is Personal: The Importance of Personalization for Recognizing Self-Reported Emotion

语音情感识别 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4215 字 阅读 →
论文解读

Joint Autoregressive Modeling of Multi-Talker Overlapped Speech Recognition and Translation

语音识别 语音翻译 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4119 字 阅读 →
论文解读

Joint Estimation of Piano Dynamics and Metrical Structure with a Multi-Task Multi-Scale Network

音乐理解 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4795 字 阅读 →
论文解读

Malefa: Multi-Granularity Learning and Effective False Alarm Suppression for Zero-Shot Keyword Spotting

零样本关键词检测 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4936 字 阅读 →
论文解读

Matrix-Structured Hierarchical Convolutional Modeling for Pronunciation Assessment and Mispronunciation Detection

语音评估 | 8.0/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6402 字 阅读 →
论文解读

MC-MRX: Reference- and Midi-Guided Music Source Extraction with Contrastive Learning

音乐源提取 | 7.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5478 字 阅读 →
论文解读

Melos: Sentence-To-Section Training with Multi-Task Learning for LLM-Driven Song Generation

音乐生成 | 6.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4737 字 阅读 →
论文解读

Mixtures of Lightweight Articulatory Experts for Multilingual Asr

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3929 字 阅读 →
论文解读

ML-SAN: Multi-Level Speaker-Adaptive Network for Emotion Recognition in Conversations

语音情感识别 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4331 字 阅读 →
论文解读

MNV-17: A High-Quality Performative Mandarin Dataset for Nonverbal Vocalization Recognition in Speech

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3597 字 阅读 →
论文解读

MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-Token Prediction

语音翻译 | 8.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5619 字 阅读 →
论文解读

Multi-Task Learning For Speech Quality Assessment Using ASR-Derived Entropy Features

语音质量评估 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4836 字 阅读 →
论文解读

Multi-Task Transformer for Explainable Speech Deepfake Detection via Formant Modeling

语音伪造检测 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4283 字 阅读 →
论文解读

NeuroSIFT: A Biologically-Inspired Framework with Explicit Signal-Noise Separation for Robust Multimodal Emotion Recognition

多模态情感识别 | 8.0/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3891 字 阅读 →
论文解读

Obstructive Sleep Apnea Endotype Prediction During Wakefulness Using Voice Biomarkers

语音生物标志物 | 6.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3906 字 阅读 →
论文解读

OMNI-AVSR: Towards Unified Multimodal Speech Recognition With Large Language Models

语音识别 | 8.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4649 字 阅读 →
论文解读

One Model–Three Tasks: Discovering a Shared Winning Ticket for Low-Complexity Audio Intelligence

音频分类 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3730 字 阅读 →