论文解读

Prompt-Guided Mixture-of-Experts for Robust Multimodal Sentiment Analysis with Missing Modalities

语音情感识别 | 8.5/10

 · 更新于 2026-09-16 · 约 10 分钟 · 4723 字 阅读 →
论文解读

PromptSep: Generative Audio Separation Via Multimodal Prompting

语音分离 | 7.5/10

 · 更新于 2026-09-16 · 约 8 分钟 · 3891 字 阅读 →
论文解读

Prosody-Guided Harmonic Attention for Phase-Coherent Neural Vocoding in the Complex Spectrum

语音合成 | 8.0/10

 · 更新于 2026-09-16 · 约 10 分钟 · 4866 字 阅读 →
论文解读

PROST-LLM: Progressively Enhancing the Speech-to-Speech Translation Capability in LLMs

语音翻译 | 7.5/10

 · 更新于 2026-09-16 · 约 12 分钟 · 5541 字 阅读 →
论文解读

Prototype-Guided Cross-Modal Contrastive Learning for Continual Audio-Visual Sound Separation

语音分离 | 7.5/10

 · 更新于 2026-09-16 · 约 9 分钟 · 4252 字 阅读 →
论文解读

PRSA: Preventing Malicious Speaker Recognition and Speech Synthesis Simultaneously with Adversarial Examples

语音匿名化 | 7.0/10

 · 更新于 2026-09-16 · 约 9 分钟 · 4445 字 阅读 →
论文解读

PSP: An Interpretable Per-Dimension Accent Benchmark for Indic Text-to-Speech

基准测试 | 7.5/10

 · 更新于 2026-09-16 · 约 11 分钟 · 5099 字 阅读 →
论文解读

PSTalker: Realistic 3D Talking Head Synthesis via a Semantic-Aware Audio-Driven Point-Based Shape

说话人合成 | 7.5/10

 · 更新于 2026-09-16 · 约 9 分钟 · 4433 字 阅读 →
论文解读

Purification Before Fusion: Toward Mask-Free Speech Enhancement for Robust Audio-Visual Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-16 · 约 9 分钟 · 4265 字 阅读 →
论文解读

Qastanet: A DNN-Based Quality Metric for Spatial Audio

空间音频 | 7.5/10

 · 更新于 2026-09-16 · 约 10 分钟 · 4668 字 阅读 →
论文解读

QE-XVC: Zero-Shot Cross-Lingual Voice Conversion via Query-Enhancement and Conditional Flow Matching

语音转换 | 7.5/10

 · 更新于 2026-09-16 · 约 10 分钟 · 4561 字 阅读 →
论文解读

QFOCUS: Controllable Synthesis for Automated Speech Stress Editing to Deliver Human-Like Emphatic Intent

语音合成 | 7.5/10

 · 更新于 2026-09-16 · 约 4 分钟 · 1766 字 阅读 →
论文解读

Quality Assessment of Noisy and Enhanced Speech with Limited Data: UWB-NTIS System for Voicemos 2024

语音质量评估 | 7.0/10

 · 更新于 2026-09-16 · 约 11 分钟 · 5174 字 阅读 →
论文解读

Quantifying Speaker Embedding Phonological Rule Interactions in Accented Speech Synthesis

语音合成 | 7.0/10

 · 更新于 2026-09-16 · 约 9 分钟 · 4134 字 阅读 →
论文解读

Random Matrix-Driven Graph Representation Learning For Bioacoustic Recognition

生物声学 | 7.5/10

 · 更新于 2026-09-16 · 约 9 分钟 · 4382 字 阅读 →
论文解读

Ranking The Impact of Contextual Specialization in Neural Speech Enhancement

语音增强 | 7.5/10

 · 更新于 2026-09-16 · 约 9 分钟 · 4252 字 阅读 →
论文解读

RAP: Real-Time Audio-Driven Portrait Animation with Video Diffusion Transformer

音视频 | 7.0/10

 · 更新于 2026-09-16 · 约 11 分钟 · 5050 字 阅读 →
论文解读

RAS: a Reliability Oriented Metric for Automatic Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-16 · 约 9 分钟 · 4136 字 阅读 →
论文解读

RASD-SR: A Robust Anomalous Sound Detection Framework with Score Recalibration

异常声音检测 | 8.5/10

 · 更新于 2026-09-16 · 约 11 分钟 · 5155 字 阅读 →
论文解读

Rationale-Guided Learning for Multimodal Emotion Recognition

语音情感识别 | 7.0/10

 · 更新于 2026-09-16 · 约 10 分钟 · 4768 字 阅读 →