论文解读

Mitigating Intra-Speaker Variability in Diarization with Style-Controllable Speech Augmentation

说话人日志 | 7.0/10

 · 更新于 2026-09-16 · 约 9 分钟 · 4023 字 阅读 →
论文解读

Mitigating Language Prior-Induced Hallucinations via Bi-Level Contrastive Decoding

多模态模型 | 7.5/10

 · 更新于 2026-09-16 · 约 8 分钟 · 3934 字 阅读 →
论文解读

Mitigating Shared-Private Branch Imbalance via Dual-Branch Rebalancing for Multimodal Sentiment Analysis

多模态模型 | 7.5/10

 · 更新于 2026-09-16 · 约 11 分钟 · 5421 字 阅读 →
论文解读

用带噪混合当高噪声步监督:Mix2Morph 如何把加法叠加训练成可控的声音注入

📄 用带噪混合当高噪声步监督:Mix2Morph 如何把加法叠加训练成可控的声音注入 会议论文 ID:conference:icassp:2026:icassp-arnumber:11460386

 · 更新于 2026-09-16 · 约 16 分钟 · 7807 字 阅读 →
论文解读

MixGAN-based Non-blind Bandwidth Extension for Audio Codec

音频增强 | 8.0/10

 · 更新于 2026-09-16 · 约 10 分钟 · 4791 字 阅读 →
论文解读

Mixture of Experts for Recognizing Depression from Interview and Reading Tasks

语音生物标志物 | 6.0/10

 · 更新于 2026-09-16 · 约 11 分钟 · 5201 字 阅读 →
论文解读

Mixture To Beamformed Mixture: Leveraging Beamformed Mixture As Weak-Supervision for Speech Enhancement and Noise-Robust ASR

语音增强 | 8.0/10

 · 更新于 2026-09-16 · 约 9 分钟 · 4195 字 阅读 →
论文解读

Mixture-of-Experts Based Soft-Label Learning for Multi-Label Speech Emotion Recognition

语音情感识别 | 7.5/10

 · 更新于 2026-09-16 · 约 10 分钟 · 4604 字 阅读 →
论文解读

Mixture-of-Experts Framework for Field-of-View Enhanced Signal-Dependent Binauralization of Moving Talkers

空间音频 | 6.5/10

 · 更新于 2026-09-16 · 约 8 分钟 · 3871 字 阅读 →
论文解读

Mixtures of Lightweight Articulatory Experts for Multilingual Asr

语音识别 | 7.0/10

 · 更新于 2026-09-16 · 约 8 分钟 · 3929 字 阅读 →
论文解读

ML-SAN: Multi-Level Speaker-Adaptive Network for Emotion Recognition in Conversations

语音情感识别 | 8.0/10

 · 更新于 2026-09-16 · 约 9 分钟 · 4331 字 阅读 →
论文解读

MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation

语音分离 | 8.0/10

 · 更新于 2026-09-16 · 约 10 分钟 · 4749 字 阅读 →
论文解读

MMEB-V3: Measuring the Performance Gaps of Omni-Modality Embedding Models

基准测试 | 7.5/10

 · 更新于 2026-09-16 · 约 9 分钟 · 4488 字 阅读 →
论文解读

MNV-17: A High-Quality Performative Mandarin Dataset for Nonverbal Vocalization Recognition in Speech

语音识别 | 7.5/10

 · 更新于 2026-09-16 · 约 8 分钟 · 3597 字 阅读 →
论文解读

Modeling Both Intra- And Inter-Utterance Variability for Conversational Emotion Recognition

语音情感识别 | 6.5/10

 · 更新于 2026-09-16 · 约 9 分钟 · 4447 字 阅读 →
论文解读

Modeling Inter-Segment Relationships in Speech for Dementia Detection with Audio Spectrogram Transformers and Graph Attention Networks

语音生物标志物 | 7.0/10

 · 更新于 2026-09-16 · 约 9 分钟 · 4152 字 阅读 →
论文解读

Modeling Strategies For Speech Enhancement in The Latent Space of a Neural Audio Codec

语音增强 | 8.0/10

 · 更新于 2026-09-16 · 约 10 分钟 · 4963 字 阅读 →
论文解读

Monitoring exposure-length variations in submarine power cables using distributed fiber-optic sensing

音频事件检测 | 6.5/10

 · 更新于 2026-09-16 · 约 7 分钟 · 3073 字 阅读 →
论文解读

More Than a Shortcut: A Hyperbolic Approach to Early-Exit Networks

音频事件检测 | 8.0/10

 · 更新于 2026-09-16 · 约 11 分钟 · 5144 字 阅读 →
论文解读

Motionbeat: Motion-Aligned Music Representation via Embodied Contrastive Learning and Bar-Equivariant Contact-Aware Encoding

舞蹈生成 | 7.5/10

 · 更新于 2026-09-16 · 约 9 分钟 · 4092 字 阅读 →