论文解读

Physics-Informed Neural Networks for Ocean Acoustic Field Reconstruction and Source Localization

声源定位 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3605 字 阅读 →
论文解读

Pianoroll-Event: A Novel Score Representation for Symbolic Music

音乐生成 | 6.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4100 字 阅读 →
论文解读

PICOAUDIO2: Temporal Controllable Text-to-Audio Generation with Natural Language Description

音频生成 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4651 字 阅读 →
论文解读

Plug-and-Play Emotion Graphs for Compositional Prompting in Zero-Shot Speech Emotion Recognition

语音情感识别 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4538 字 阅读 →
论文解读

Poly-SVC: Polyphony-Aware Singing Voice Conversion with Harmonic Modeling

歌唱语音转换 | 6.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4949 字 阅读 →
论文解读

Polynomial Mixing for Efficient Self-Supervised Speech Encoders

语音识别 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4653 字 阅读 →
论文解读

Position-Invariant Fine-Tuning Of Speech Enhancement Models With Self-Supervised Speech Representations

语音增强 | 6.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4224 字 阅读 →
论文解读

Principled Coarse-Grained Acceptance For Speculative Decoding In Speech

语音合成 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3911 字 阅读 →
论文解读

PRoADS: Provably Secure And Robust Audio Diffusion Steganography With Latent Optimization And Backward Euler Inversion

音频安全 | 6.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3627 字 阅读 →
论文解读

Probing the Hidden Talent of ASR foundation models for L2 English Oral Assessment

预训练 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4262 字 阅读 →
论文解读

Probing Whisper for Dysarthric Speech in Detection and Assessment

语音生物标志物 | 6.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3926 字 阅读 →
论文解读

Production-Scale Dynamic Vocabulary ASR Biasing with Word-Level FST and Robust Training

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3562 字 阅读 →
论文解读

Proficiency-Aware Adaptation and Data Augmentation for Robust L2 ASR

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3992 字 阅读 →
论文解读

Prompt-Guided Mixture-of-Experts for Robust Multimodal Sentiment Analysis with Missing Modalities

语音情感识别 | 8.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4723 字 阅读 →
论文解读

PromptSep: Generative Audio Separation Via Multimodal Prompting

语音分离 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3891 字 阅读 →
论文解读

Prosody-Guided Harmonic Attention for Phase-Coherent Neural Vocoding in the Complex Spectrum

语音合成 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4866 字 阅读 →
论文解读

PROST-LLM: Progressively Enhancing the Speech-to-Speech Translation Capability in LLMs

语音翻译 | 7.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5541 字 阅读 →
论文解读

Prototype-Guided Cross-Modal Contrastive Learning for Continual Audio-Visual Sound Separation

语音分离 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4252 字 阅读 →
论文解读

PRSA: Preventing Malicious Speaker Recognition and Speech Synthesis Simultaneously with Adversarial Examples

语音匿名化 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4445 字 阅读 →
论文解读

PSTalker: Realistic 3D Talking Head Synthesis via a Semantic-Aware Audio-Driven Point-Based Shape

说话人合成 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4433 字 阅读 →