论文解读

Triage Knowledge Distillation for Speaker Verification

说话人验证 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4613 字 阅读 →
论文解读

TTA: Transcribe, Translate and Alignment for Cross-Lingual Speech Representation

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5261 字 阅读 →
论文解读

TVP-UNet: Threshold Variance Penalty U-Net for Voice Activity Detection in Dysarthric Speech

语音活动检测 | 7.0/10

 · 更新于 2026-09-06 · 约 7 分钟 · 3256 字 阅读 →
论文解读

Two-Stage Language Model Framework for Acoustic Echo Cancellation

语音增强 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4401 字 阅读 →
论文解读

UJCodec: An End-to-end Unet-Style Codec for Joint Speech Compression and Enhancement

语音增强 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4672 字 阅读 →
论文解读

UMA-SPLIT: Unimodal Aggregation for Both English and Mandarin Non-Autoregressive Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3808 字 阅读 →
论文解读

UMV: A Mixture-Of-Experts Vision Transformer with Multi-Spectrogram Fusion for Underwater Ship Noise Classification

音频分类 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4212 字 阅读 →
论文解读

Uncertainty-Aware 3D Emotional Talking Face Synthesis with Emotion Prior Distillation

音视频 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4803 字 阅读 →
论文解读

Understanding Textual Capability Degradation in Speech LLMS via Parameter Importance Analysis

语音问答 | 7.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5549 字 阅读 →
论文解读

Understanding the Strengths and Weaknesses of SSL Models for Audio Deepfake Model Attribution

音频深度伪造检测 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4679 字 阅读 →
论文解读

Universr: Unified and Versatile Audio Super-Resolution Via Vocoder-Free Flow Matching

音频超分辨率 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4168 字 阅读 →
论文解读

UNMIXX: Untangling Highly Correlated Singing Voices Mixtures

语音分离 | 8.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4383 字 阅读 →
论文解读

Unseen but Not Unknown: Using Dataset Concealment to Robustly Evaluate Speech Quality Estimation Models

语音质量评估 | 8.3/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4293 字 阅读 →
论文解读

Unsupervised Discovery and Analysis of the Vocal Repertoires and Patterns of Select Corvid Species

生物声学 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4918 字 阅读 →
论文解读

Unsupervised Lexicon Learning from Speech is Limited by Representations Rather than Clustering

语音发现 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4486 字 阅读 →
论文解读

USVexplorer: Robust Detection of Ultrasonic Vocalizations with Cross Species Generalization

音频事件检测 | 8.0/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5686 字 阅读 →
论文解读

UTI-LLM: A Personalized Articulatory-Speech Therapy Assistance System Based on Multimodal Large Language Model

语音对话系统 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4909 字 阅读 →
论文解读

Utilizing Information Theoretic Approach to Study Cochlear Neural Degeneration

生物声学 | 6.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3850 字 阅读 →
论文解读

UVT-LM: Unifying Visual and Tactile Perception with Language Model

跨模态 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4823 字 阅读 →
论文解读

V2A-DPO: Omni-Preference Optimization for Video-To-Audio Generation

视频到音频生成 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4613 字 阅读 →