论文解读

SCOPE: Entanglement Frontier Escape for Source-Free Class Unlearning

说话人验证 | 6.3/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6134 字 阅读 →
论文解读

Extracting Voice Styles from Frozen TTS Models via Gradient-Based Inverse Optimization

语音克隆 | 7.9/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5923 字 阅读 →
论文解读

Simple Language Normalization Wins: Cross-Lingual Speaker Verification for the TidyVoice 2026 Challenge

说话人验证 | 6.8/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6547 字 阅读 →
论文解读

Improving the performance of an ASV system using hybrid speech features

说话人验证 | 5.0/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6562 字 阅读 →
论文解读

Multimodal Speaker Verification as a Threat to Speaker Anonymization

说话人验证 | 9.2/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7725 字 阅读 →
论文解读

AMECxSV: Adaptive Metadata-Driven Embedding-Fusion Calibration for X-Lingual Speaker Verification

说话人验证 | 7.8/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7571 字 阅读 →
论文解读

Do Speech Tokens Leak Voiceprints? Speaker Inversion Attacks Against End-to-End Speech Language Models

说话人验证 | 7.7/10

 · 更新于 2026-09-24 · 约 17 分钟 · 8389 字 阅读 →
论文解读

A Geometry-Limited Identification Floor and Its Consequences for Voice-Clone Attribution in Professional Voice Actors

说话人验证 | 8.8/10

 · 更新于 2026-09-24 · 约 17 分钟 · 8183 字 阅读 →
论文解读

Rethinking Speech Foundation Model Fine-tuning: Better SFT or Better Match?

语音情感识别 | 6.7/10

 · 更新于 2026-09-24 · 约 18 分钟 · 8730 字 阅读 →
论文解读

Listen to the Features: Voice Anonymization Driven by Content Embedding Matching over Signal Reconstruction

语音克隆 | 6.5/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6792 字 阅读 →
论文解读

Text-Independent Speaker Verification Using Discrete Audio Tokens

说话人验证 | 5.2/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7091 字 阅读 →
论文解读

Flow Matching-Based Speech Source Separation with Best-of-N Biometric Sampling

语音分离 | 4.9/10

 · 更新于 2026-09-24 · 约 17 分钟 · 8227 字 阅读 →
论文解读

Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment

语音情感识别 | 5.5/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7228 字 阅读 →
论文解读

ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions

语音合成 | 7.6/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6226 字 阅读 →
论文解读

Streaming Neural Speech Codecs through Time-Invariant Representations

语音编码 | 6.0/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7723 字 阅读 →
论文解读

DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video Generation

音视频生成 | 8/10

 · 更新于 2026-09-24 · 约 20 分钟 · 9827 字 阅读 →
论文解读

Disentangling Speaker and Language Effects in Cross-Lingual Speaker Verification for Iberian Languages

说话人验证 | 5.6/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5518 字 阅读 →
论文解读

Kiwano: A Cutting-Edge Open-Source Toolkit for Speaker Verification

说话人验证 | 7.6/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4879 字 阅读 →
论文解读

LISE : Listenable Interpretable Speaker Embeddings

说话人验证 | 6.8/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5422 字 阅读 →
论文解读

Speaker Identity in Non-Verbal Vocalizations: Conditional Distillation and Mixture of Experts Approach

说话人验证 | 9.1/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6090 字 阅读 →