论文解读

Uncertainty-Aware 3D Emotional Talking Face Synthesis with Emotion Prior Distillation

音视频 | 8.0/10

 · 更新于 2026-09-16 · 约 10 分钟 · 4803 字 阅读 →
论文解读

Understanding Textual Capability Degradation in Speech LLMS via Parameter Importance Analysis

语音问答 | 7.5/10

 · 更新于 2026-09-16 · 约 12 分钟 · 5549 字 阅读 →
论文解读

Understanding the Strengths and Weaknesses of SSL Models for Audio Deepfake Model Attribution

音频深度伪造检测 | 7.0/10

 · 更新于 2026-09-16 · 约 10 分钟 · 4679 字 阅读 →
论文解读

UNet-Based Fusion and Exponential Moving Average Adaptation for Noise-Robust Speaker Recognition

说话人验证 | 7.5/10

 · 更新于 2026-09-16 · 约 11 分钟 · 5394 字 阅读 →
论文解读

Universr: Unified and Versatile Audio Super-Resolution Via Vocoder-Free Flow Matching

音频超分辨率 | 8.0/10

 · 更新于 2026-09-16 · 约 9 分钟 · 4168 字 阅读 →
论文解读

UNMIXX: Untangling Highly Correlated Singing Voices Mixtures

语音分离 | 8.5/10

 · 更新于 2026-09-16 · 约 9 分钟 · 4383 字 阅读 →
论文解读

Unrequited Emotions: Investigating the Gaps in Motivation and Practice in Speech Emotion Recognition Research

语音情感识别 | 8.0/10

 · 更新于 2026-09-16 · 约 6 分钟 · 2880 字 阅读 →
论文解读

Unseen but Not Unknown: Using Dataset Concealment to Robustly Evaluate Speech Quality Estimation Models

语音质量评估 | 8.3/10

 · 更新于 2026-09-16 · 约 9 分钟 · 4293 字 阅读 →
论文解读

Unsupervised Discovery and Analysis of the Vocal Repertoires and Patterns of Select Corvid Species

生物声学 | 7.5/10

 · 更新于 2026-09-16 · 约 10 分钟 · 4918 字 阅读 →
论文解读

Unsupervised Lexicon Learning from Speech is Limited by Representations Rather than Clustering

语音发现 | 8.0/10

 · 更新于 2026-09-16 · 约 9 分钟 · 4486 字 阅读 →
论文解读

USVexplorer: Robust Detection of Ultrasonic Vocalizations with Cross Species Generalization

音频事件检测 | 8.0/10

 · 更新于 2026-09-16 · 约 12 分钟 · 5686 字 阅读 →
论文解读

UTI-LLM: A Personalized Articulatory-Speech Therapy Assistance System Based on Multimodal Large Language Model

语音对话系统 | 7.5/10

 · 更新于 2026-09-16 · 约 10 分钟 · 4909 字 阅读 →
论文解读

Utilizing Information Theoretic Approach to Study Cochlear Neural Degeneration

生物声学 | 6.5/10

 · 更新于 2026-09-16 · 约 8 分钟 · 3850 字 阅读 →
论文解读

UVT-LM: Unifying Visual and Tactile Perception with Language Model

跨模态 | 7.0/10

 · 更新于 2026-09-16 · 约 10 分钟 · 4823 字 阅读 →
论文解读

V2A-DPO: Omni-Preference Optimization for Video-To-Audio Generation

视频到音频生成 | 7.5/10

 · 更新于 2026-09-16 · 约 10 分钟 · 4613 字 阅读 →
论文解读

Variational Low-Rank Adaptation for Personalized Impaired Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-16 · 约 11 分钟 · 5297 字 阅读 →
论文解读

VBx for End-to-End Neural and Clustering-Based Diarization

说话人分离 | 8.5/10

 · 更新于 2026-09-16 · 约 10 分钟 · 4816 字 阅读 →
论文解读

VChangeCodec: An Ultra Low-Complexity Neural Speech Codec with Built-In Voice Changer for Customized Real-Time Communication

语音转换 语音增强 | 8.0/10

 · 更新于 2026-09-16 · 约 11 分钟 · 5237 字 阅读 →
论文解读

Via Score to Performance: Efficient Human-Controllable Long Song Generation with Bar-Level Symbolic Notation

音乐生成 | 7.5/10

 · 更新于 2026-09-16 · 约 10 分钟 · 4896 字 阅读 →
论文解读

Vib2Sound: Separation Of Multimodal Sound Sources

语音分离 | 6.5/10

 · 更新于 2026-09-16 · 约 10 分钟 · 4522 字 阅读 →