论文解读

Decoder-Only Conformer with Modality-Aware Sparse Mixtures of Experts for ASR

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5265 字 阅读 →
论文解读

DMP-TTS: Disentangled Multi-Modal Prompting for Controllable Text-to-Speech with Chained Guidance

语音合成 | 7.5/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6161 字 阅读 →
论文解读

DPO-Regularized Regression for Age Prediction

说话人识别 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4470 字 阅读 →
论文解读

DSpAST: Disentangled Representations for Spatial Audio Reasoning with Large Language Models

音频问答 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4732 字 阅读 →
论文解读

Dual Data Scaling for Robust Two-Stage User-Defined Keyword Spotting

语音活动检测 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5107 字 阅读 →
论文解读

Dual-Strategy-Enhanced Conbimamba for Neural Speaker Diarization

说话人分离 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4642 字 阅读 →
论文解读

Dynamic Balanced Cross-Modal Attention with Gated Sequence Restoration: Towards Robust Multimodal Sentiment Analysis

跨模态 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4284 字 阅读 →
论文解读

E2E-AEC: Implementing An End-To-End Neural Network Learning Approach for Acoustic Echo Cancellation

语音增强 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4163 字 阅读 →
论文解读

EEG and Eye-Tracking Driven Dynamic Target Speaker Extraction with Spontaneous Attention Switching

语音分离 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4652 字 阅读 →
论文解读

EMG-to-Speech with Fewer Channels

语音合成 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4171 字 阅读 →
论文解读

EmoTri-RL: Emotion- and Cause-Aware Reinforcement Learning for Multi-Modal Empathetic Dialogue

语音情感识别 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4826 字 阅读 →
论文解读

Enhancing Speech Intelligibility Prediction for Hearing Aids with Complementary Speech Foundation Model Representations

语音增强 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3521 字 阅读 →
论文解读

Estimating Respiratory Effort from Nocturnal Breathing Sounds for Obstructive Sleep Apnoea Screening

音频分类 | 6.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3681 字 阅读 →
论文解读

From Contrast to Commonality: Audio Commonality Captioning for Enhanced Audio-Text Cross-Modal Understanding in Multimodal LLMS

音频场景理解 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4037 字 阅读 →
论文解读

From Diet to Free Lunch: Estimating Auxiliary Signal Properties Using Dynamic Pruning Masks in Speech Enhancement Networks

语音增强 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5029 字 阅读 →
论文解读

FUSEMOS: Perceptual Evaluation of Text-to-Music Generation with Dual-Encoder Fusion and Ranking-Aware Composite Loss

音乐生成 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4701 字 阅读 →
论文解读

Fusion of Multimodal Estimations by Extended State Hidden Markov Model: Application to Fetal Heart Rate Monitoring

生物声学 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4013 字 阅读 →
论文解读

GLUE: Gradient-free Learning to Unify Experts

迁移学习 | 6.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5139 字 阅读 →
论文解读

GRNet: Graph Reconstruction Network for Robust Multimodal Sentiment Analysis

多模态情感分析 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4441 字 阅读 →
论文解读

Hierarchical Activity Recognition and Captioning from Long-Form Audio

音频事件检测 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4494 字 阅读 →