论文解读

CodeSep: Low-Bitrate Codec-Driven Speech Separation with Base-Token Disentanglement and Auxiliary-Token Serial Prediction

语音分离 | 7.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6108 字 阅读 →
论文解读

CompSpoof: A Dataset and Joint Learning Framework for Component-Level Audio Anti-Spoofing Countermeasures

音频深度伪造检测 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5184 字 阅读 →
论文解读

Context-Aware Dynamic Graph Learning for Multimodal Emotion Recognition with Missing Modalities

语音情感识别 | 8.8/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4275 字 阅读 →
论文解读

Contextual Biasing for ASR in Speech LLM with Common Word Cues and Bias Word Position Prediction

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5305 字 阅读 →
论文解读

Cross-Modal Knowledge Distillation for Speech Large Language Models

语音大模型 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3931 字 阅读 →
论文解读

Decoder-Only Conformer with Modality-Aware Sparse Mixtures of Experts for ASR

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5265 字 阅读 →
论文解读

DMP-TTS: Disentangled Multi-Modal Prompting for Controllable Text-to-Speech with Chained Guidance

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6161 字 阅读 →
论文解读

DPO-Regularized Regression for Age Prediction

说话人识别 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4470 字 阅读 →
论文解读

DSpAST: Disentangled Representations for Spatial Audio Reasoning with Large Language Models

音频问答 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4732 字 阅读 →
论文解读

Dual Data Scaling for Robust Two-Stage User-Defined Keyword Spotting

语音活动检测 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5107 字 阅读 →
论文解读

Dual-Strategy-Enhanced Conbimamba for Neural Speaker Diarization

说话人分离 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4642 字 阅读 →
论文解读

E2E-AEC: Implementing An End-To-End Neural Network Learning Approach for Acoustic Echo Cancellation

语音增强 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4163 字 阅读 →
论文解读

EEG and Eye-Tracking Driven Dynamic Target Speaker Extraction with Spontaneous Attention Switching

语音分离 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4652 字 阅读 →
论文解读

EMG-to-Speech with Fewer Channels

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4171 字 阅读 →
论文解读

EmoTri-RL: Emotion- and Cause-Aware Reinforcement Learning for Multi-Modal Empathetic Dialogue

语音情感识别 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4826 字 阅读 →
论文解读

Enhancing Speech Intelligibility Prediction for Hearing Aids with Complementary Speech Foundation Model Representations

语音增强 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3521 字 阅读 →
论文解读

Estimating Respiratory Effort from Nocturnal Breathing Sounds for Obstructive Sleep Apnoea Screening

音频分类 | 6.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3681 字 阅读 →
论文解读

From Contrast to Commonality: Audio Commonality Captioning for Enhanced Audio-Text Cross-Modal Understanding in Multimodal LLMS

音频场景理解 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4037 字 阅读 →
论文解读

From Diet to Free Lunch: Estimating Auxiliary Signal Properties Using Dynamic Pruning Masks in Speech Enhancement Networks

语音增强 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5029 字 阅读 →
论文解读

FUSEMOS: Perceptual Evaluation of Text-to-Music Generation with Dual-Encoder Fusion and Ranking-Aware Composite Loss

音乐生成 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4701 字 阅读 →