论文解读

Transcript-Free Flow-Matching Text-to-Speech via Speech Feature Conditioning

语音合成 | 7.7/10

 · 更新于 2026-09-07 · 约 10 分钟 · 4674 字 阅读 →
论文解读

Zero-VC: Zero-Lookahead Streaming Voice Conversion via Speaker Anonymization

语音转换 | 6.1/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5580 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-19

共分析 40 篇语音/AI 论文

 · 更新于 2026-09-07 · 约 110 分钟 · 54842 字 阅读 →
论文解读

A Survey of Methods for the Discretization of Phonograph Record Playback Filters

A Survey of Methods for the Discretization of Phonograph Record Playback Filters

 · 更新于 2026-09-07 · 约 11 分钟 · 5415 字 阅读 →
论文解读

Adaptive Speech-to-Spike Encoding for Spiking Neural Networks

Adaptive Speech-to-Spike Encoding for Spiking Neural Networks

 · 更新于 2026-09-07 · 约 11 分钟 · 5362 字 阅读 →
论文解读

Audio-to-Audio via Diffusion Warm Initialization

音频生成 | 7.6/10

 · 更新于 2026-09-07 · 约 13 分钟 · 6293 字 阅读 →
论文解读

Augmenting Dysarthric Speech Severity Assessment with MOS Supervision

语音质量评估 | 7/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5497 字 阅读 →
论文解读

Beyond AHI: An Interpretable Causal-Discovery-Guided Framework for Sleep Recovery in Connected Health

Beyond AHI: An Interpretable Causal-Discovery-Guided Framework for Sleep Recovery in Connected Health

 · 更新于 2026-09-07 · 约 11 分钟 · 5122 字 阅读 →
论文解读

Closing the Loop: PID Feedback Control for Interpretable Activation Steering in Symbolic Music Generation

音乐生成 | 8.7/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5239 字 阅读 →
论文解读

Constraining to Generalize: Subspace Tuning for Few-shot Generalization of Audio-Language Models

音频分类 | 7.5/10

 · 更新于 2026-09-07 · 约 16 分钟 · 7536 字 阅读 →
论文解读

Continuous Audio Thinking for Large Audio Language Models

Continuous Audio Thinking for Large Audio Language Models

 · 更新于 2026-09-07 · 约 25 分钟 · 12070 字 阅读 →
论文解读

Continuous-Speech Parkinson's Disease Detection Using Acoustic and Inharmonicity Features

Continuous-Speech Parkinson's Disease Detection Using Acoustic and Inharmonicity Features

 · 更新于 2026-09-07 · 约 14 分钟 · 6751 字 阅读 →
论文解读

DASH: Dual-View Self-Distillation with Multi-Layer Hidden Representations for Robust Speech Recognition

语音识别 | 6.6/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5965 字 阅读 →
论文解读

EMORSION: Examining the Impact of Audio Parameters on Emotional Responses and Immersion in Film

EMORSION: Examining the Impact of Audio Parameters on Emotional Responses and Immersion in Film

 · 更新于 2026-09-07 · 约 13 分钟 · 6057 字 阅读 →
论文解读

Evaluating Dynamic Range Compressor Models Using Control-Voltage Measurements: an Approach and Dataset

模型评估 | 7.8/10

 · 更新于 2026-09-07 · 约 10 分钟 · 5005 字 阅读 →
论文解读

Fair Cognitive Impairment Detection Through Unlearning

多模态模型 | 7.7/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5572 字 阅读 →
论文解读

FineCombo-TTS: Collaborative and Precise Controllable Speech Synthesis Using Text Descriptions and Reference Speech

语音合成 | 7.6/10

 · 更新于 2026-09-07 · 约 23 分钟 · 11197 字 阅读 →
论文解读

Generalised Transcoding Framework for Arbitrary Spatial Audio Capture and Playback Formats

空间音频 | 8.7/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5840 字 阅读 →
论文解读

GRIDEX: Grid-Grounded Forensic Explanations for Deepfake Spectrogram Analysis

语音合成 | 8.6/10

 · 更新于 2026-09-07 · 约 16 分钟 · 7676 字 阅读 →
论文解读

Human-AI Coevolution Dynamics: A Formal Theory of Social Intelligence Emergence Through Long-Term Interaction

Human-AI Coevolution Dynamics: A Formal Theory of Social Intelligence Emergence Through Long-Term Interaction

 · 更新于 2026-09-07 · 约 10 分钟 · 4617 字 阅读 →