论文解读

MT-HuBERT: Self-Supervised Mix-Training for Few-Shot Keyword Spotting in Mixed Speech

关键词检测 | 7.0/10

 · 更新于 2026-09-06 · 约 23 分钟 · 11472 字 阅读 →
论文解读

MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-Token Prediction

语音翻译 | 8.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5619 字 阅读 →
论文解读

Multi-Channel Speech Enhancement for Cocktail Party Speech Emotion Recognition

语音情感识别 | 7.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5585 字 阅读 →
论文解读

Multi-Layer Attentive Probing Improves Transfer of Audio Representations for Bioacoustics

生物声学 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4488 字 阅读 →
论文解读

Multi-Scale Physiologically-Motivated Alignment for Auditory Attention Decoding

听觉注意力解码 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4347 字 阅读 →
论文解读

Multi-Task Learning For Speech Quality Assessment Using ASR-Derived Entropy Features

语音质量评估 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4836 字 阅读 →
论文解读

Multi-Task Transformer for Explainable Speech Deepfake Detection via Formant Modeling

语音伪造检测 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4283 字 阅读 →
论文解读

Multi-View Hierarchical Hypergraph Neural Network for Automatic Stuttering Detection

语音生物标志物 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5262 字 阅读 →
论文解读

Multilingual Supervised Pretraining with Lm-Assisted Decoding for Visual Speech Recognition

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3573 字 阅读 →
论文解读

Multimodal Co-Training with Subtractive Unlabeled-Benefit Bounds

多模态学习 | 6.0/10

 · 更新于 2026-09-06 · 约 7 分钟 · 3241 字 阅读 →
论文解读

Multimodal Fusion-Based IPCLIP Network for Mixed Reality Surgical Assistance

多模态模型 | 6.5/10

 · 更新于 2026-09-06 · 约 7 分钟 · 3263 字 阅读 →
论文解读

Multimodal LLMs as Expert Speech Annotators: Acoustic Macro-Descriptors for Parkinson's Detection

语音生物标志物 | 6.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4079 字 阅读 →
论文解读

Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching

音频生成 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5404 字 阅读 →
论文解读

Multimodal Self-Attention Network with Temporal Alignment for Audio-Visual Emotion Recognition

语音情感识别 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4714 字 阅读 →
论文解读

Multimodal Transformer with Multiperspective Training for Predicting Self-Expression Skills from Video Interview

多模态模型 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4288 字 阅读 →
论文解读

Multimodal Variational Graph Network for Multimodal Sentiment Analysis

语音情感识别 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5407 字 阅读 →
论文解读

MuseTok: Symbolic Music Tokenization for Generation and Semantic Understanding

音乐生成 | 8.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5119 字 阅读 →
论文解读

Musicdetr: A Position-Aware Spectral Note Detection Model for Singing Transcription

歌唱语音转录 | 8.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4470 字 阅读 →
论文解读

MusiCRS: Benchmarking Audio-Centric Conversational Recommendation

音乐推荐 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4258 字 阅读 →
论文解读

Natural Language to Spatial Audio Parameters: Lightweight Deterministic Rendering for Creative Authoring

空间音频 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4737 字 阅读 →