论文解读

ProsoCodec: Prosody-Oriented Speech Codec for Voice Conversion

语音转换 | 6.6/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5546 字 阅读 →
论文解读

Scaling Audio Models Efficiently: A Joint Study of Compute Constraints and Optimization Behavior

语音识别 | 7.2/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5153 字 阅读 →
论文解读

SDP-Codec: A Speaker-Decoupled Speech Codec with Pitch Injection for Low-Bitrate Coding and Zero-Shot Voice Conversion

语音编码 | 7.2/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5887 字 阅读 →
论文解读

Sea-Scan: High-Accuracy, ML-based Dark Vessel Detection and Localisation via Weakly Supervised DAS Monitoring

Sea-Scan: High-Accuracy, ML-based Dark Vessel Detection and Localisation via Weakly Supervised DAS Monitoring

 · 更新于 2026-09-07 · 约 19 分钟 · 9441 字 阅读 →
论文解读

Speaker Identity in Non-Verbal Vocalizations: Conditional Distillation and Mixture of Experts Approach

说话人验证 | 9.1/10

 · 更新于 2026-09-07 · 约 13 分钟 · 6090 字 阅读 →
论文解读

STAR-VAE: Structured Topology-Aware Regularization for Audio Reconstruction and Generation

音频生成 | 8.8/10

 · 更新于 2026-09-07 · 约 24 分钟 · 11842 字 阅读 →
论文解读

Streaming T5-based Text-to-Speech Synthesis with Limited Lookahead

语音合成 | 6.7/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5591 字 阅读 →
论文解读

Synthesizing the Lombard Effect: Multi-Level Control of Speech Clarity and Vocal Effort in TTS

语音合成 | 7.2/10

 · 更新于 2026-09-07 · 约 10 分钟 · 4932 字 阅读 →
论文解读

The Anatomy of the CTC Oracle Gap: Acoustic Exhaustion and Linguistic Recovery

语音识别 | 7.3/10

 · 更新于 2026-09-07 · 约 15 分钟 · 7447 字 阅读 →
论文解读

The Watermark Shortcut: How Provenance Marking Sabotages Audio Deepfake Detection

数据增强 | 6.8/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5762 字 阅读 →
论文解读

Time-Frequency Weighted Losses for Phoneme Reconstruction in DNN-Based Speech Enhancement

语音增强 | 7.8/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5965 字 阅读 →
论文解读

Toward Open-Set Speaker Attribute Prediction with Keyword-Appended LLM Embeddings

多模态模型 | 7.8/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5459 字 阅读 →
论文解读

Towards Detecting Neural Audio Codec Synthesized Heart Sounds

自监督学习 | 8.7/10

 · 更新于 2026-09-07 · 约 14 分钟 · 6772 字 阅读 →
论文解读

Unlocking In-Context Learning in Audio-Language Models from Decentralized Medical Audio

联邦学习 | 7/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5185 字 阅读 →
论文解读

Using Phonological-Level Wav2Vec2 for Mandarin Automatic Mispronunciation Detection and Diagnosis

语音识别 | 8.3/10

 · 更新于 2026-09-07 · 约 13 分钟 · 6199 字 阅读 →
论文解读

Vaani Benchmark V1.0: An Inclusive Multimodal Benchmark Dataset for Hindi

语音识别 | 6.7/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5011 字 阅读 →
论文解读

What Do Neural Networks Learn for TDOA Estimation? A Cross-Architecture Probing Study

声源定位 | 7.7/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5200 字 阅读 →
论文解读

When EER Hides Deployment Failure: Auditing Threshold Transfer and Unlabeled Score Calibration for Speech Deepfake Detectors

When EER Hides Deployment Failure: Auditing Threshold Transfer and Unlabeled Score Calibration for Speech Deepfake Detectors

 · 更新于 2026-09-07 · 约 13 分钟 · 6124 字 阅读 →
论文解读

Word Lengthening as a Function of Utterance Position: A Multi-Corpus Study

语音合成 | 8.1/10

 · 更新于 2026-09-07 · 约 10 分钟 · 4828 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-23

共分析 83 篇语音/AI 论文

 · 更新于 2026-09-07 · 约 224 分钟 · 112108 字 阅读 →