论文解读

Mitigating Intra-Speaker Variability in Diarization with Style-Controllable Speech Augmentation

说话人日志 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4023 字 阅读 →
论文解读

Mitigating Language Prior-Induced Hallucinations via Bi-Level Contrastive Decoding

多模态模型 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3934 字 阅读 →
论文解读

Mix2Morph: Learning Sound Morphing from Noisy Mixes

音频生成 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4754 字 阅读 →
论文解读

MixGAN-based Non-blind Bandwidth Extension for Audio Codec

音频增强 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4791 字 阅读 →
论文解读

Mixture of Experts for Recognizing Depression from Interview and Reading Tasks

语音生物标志物 | 6.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5201 字 阅读 →
论文解读

Mixture To Beamformed Mixture: Leveraging Beamformed Mixture As Weak-Supervision for Speech Enhancement and Noise-Robust ASR

语音增强 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4195 字 阅读 →
论文解读

Mixture-of-Experts Based Soft-Label Learning for Multi-Label Speech Emotion Recognition

语音情感识别 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4604 字 阅读 →
论文解读

Mixture-of-Experts Framework for Field-of-View Enhanced Signal-Dependent Binauralization of Moving Talkers

空间音频 | 6.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3871 字 阅读 →
论文解读

Mixtures of Lightweight Articulatory Experts for Multilingual Asr

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3929 字 阅读 →
论文解读

MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation

语音分离 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4749 字 阅读 →
论文解读

MNV-17: A High-Quality Performative Mandarin Dataset for Nonverbal Vocalization Recognition in Speech

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3597 字 阅读 →
论文解读

Modeling Both Intra- And Inter-Utterance Variability for Conversational Emotion Recognition

语音情感识别 | 6.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4447 字 阅读 →
论文解读

Modeling Inter-Segment Relationships in Speech for Dementia Detection with Audio Spectrogram Transformers and Graph Attention Networks

语音生物标志物 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4152 字 阅读 →
论文解读

Modeling Strategies For Speech Enhancement in The Latent Space of a Neural Audio Codec

语音增强 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4963 字 阅读 →
论文解读

More Than a Shortcut: A Hyperbolic Approach to Early-Exit Networks

音频事件检测 | 8.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5144 字 阅读 →
论文解读

Motionbeat: Motion-Aligned Music Representation via Embodied Contrastive Learning and Bar-Equivariant Contact-Aware Encoding

舞蹈生成 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4092 字 阅读 →
论文解读

MR-FlowDPO: Multi-Reward Direct Preference Optimization for Flow-Matching Text-to-Music Generation

音乐生成 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5357 字 阅读 →
论文解读

MSANET: Multi-Scale Semantic Aggregation Network for Brain-Assisted Speech Enhancement in Multi-Speaker Conditions

语音增强 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4444 字 阅读 →
论文解读

MSCT: Differential Cross-Modal Attention for Deepfake Detection

音频深度伪造检测 | 6.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3968 字 阅读 →
论文解读

MSF-SER: Enriching Acoustic Modeling with Multi-Granularity Semantics for Speech Emotion Recognition

语音情感识别 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5146 字 阅读 →