论文解读

OV-INSTRUCTTTS: Towards Open-Vocabulary Instruct Text-to-Speech

语音合成 | 8.0/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5743 字 阅读 →
论文解读

PAC: Pronunciation-Aware Contextualized Large Language Model-Based Automatic Speech Recognition

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4114 字 阅读 →
论文解读

PADAM: Perceptual Audio Defect Assessment Model

音频分类 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4814 字 阅读 →
论文解读

ParaGSE: Parallel Generative Speech Enhancement with Group-Vector-Quantization-Based Neural Speech Codec

语音增强 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4793 字 阅读 →
论文解读

Parametric Neural Amp Modeling with Active Learning

音频生成 | 8.0/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3918 字 阅读 →
论文解读

PC-MCL: Patient-Consistent Multi-Cycle Learning with Multi-Label Bias Correction for Respiratory Sound Classification

音频分类 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 5005 字 阅读 →
论文解读

Peeking Into the Future for Contextual Biasing

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5086 字 阅读 →
论文解读

Perceptual Loss Optimized HRTF Personalization in Spherical Harmonic Domain

空间音频 | 7.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5312 字 阅读 →
论文解读

Perceptual Quality Assessment for Stylized Talking Heads

模型评估 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4192 字 阅读 →
论文解读

PerformSinger: Multimodal Singing Voice Synthesis Leveraging Synchronized Lip Cues from Singing Performance Videos

歌唱语音合成 | 4.5/10

 · 更新于 2026-09-06 · 约 3 分钟 · 1471 字 阅读 →
论文解读

Personal Sound Zones with Flexible Bright Zone Control

空间音频 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4206 字 阅读 →
论文解读

PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models

语音对话系统 | 8.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4703 字 阅读 →
论文解读

PFluxTTS: Hybrid Flow-Matching TTS with Robust Cross-Lingual Voice Cloning and Inference-Time Model Fusion

语音合成 | 7.0/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6124 字 阅读 →
论文解读

PG-SE: Predictive Acceleration and Correction for Generative Speech Enhancement

语音增强 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5158 字 阅读 →
论文解读

Phase-Retrieval-Based Physics-Informed Neural Networks For Acoustic Magnitude Field Reconstruction

声源定位 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4181 字 阅读 →
论文解读

Phase-Space Signal Processing of Acoustic Data for Advanced Manufacturing In-Situ Monitoring

音频事件检测 | 7.0/10

 · 更新于 2026-09-06 · 约 7 分钟 · 3474 字 阅读 →
论文解读

PhoenixDSR: Phoneme-Guided and LLM-Enhanced Dysarthric Speech Recognition

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5876 字 阅读 →
论文解读

Phoneme-Level Visual Speech Recognition via Point-Visual Fusion and Language Model Reconstruction

视觉语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5224 字 阅读 →
论文解读

Phonological Tokenizer: Prosody-Aware Phonetic Token Via Multi-Objective Fine-Tuning with Differentiable K-Means

语音表示学习 | 8.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5452 字 阅读 →
论文解读

Phrased: Phrase Dictionary Biasing for Speech Translation

语音翻译 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4074 字 阅读 →