论文解读

Teaching Speech Enhancement Models to Sing: Domain Adaptation from Speech Enhancement to Singing Voice Separation

音乐源分离 | 6.7/10

 · 更新于 2026-09-06 · 约 19 分钟 · 9285 字 阅读 →
论文解读

The SonicAGI System for the REAL-TSE Challenge

语音分离 | 6.8/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8024 字 阅读 →
论文解读

Tight-Frame Reconstruction for Acoustic Intensity Estimation Using Cardioid Microphone Pairs

声源定位 | 6.8/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7054 字 阅读 →
论文解读

Transcript-Free Lightweight Detection of Alzheimer's Disease from Spontaneous Speech Using Handcrafted MFCC-Dominant Acoustic Biomarkers

语音属性识别 | 4.9/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6802 字 阅读 →
论文解读

Unified Gradient Projection: Language-Balanced Continual Learning for Multilingual Low-Resource ASR

语音识别 | 7.2/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6459 字 阅读 →
论文解读

Verifier-Guided Twelve-Tone Composition: A Generate-Verify-Repair Harness for Symbolic Music Generation

音乐生成 | 6.0/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6192 字 阅读 →
论文解读

VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion

语音伪造检测 | 8.1/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6950 字 阅读 →
论文解读

WaveNet-Style Guitar Amplifier Model Pruning for Real-Time iOS Deployment

音频生成 | 8.0/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5705 字 阅读 →
论文解读

What You Train Is What You Get: Gender Bias, Training Composition, and Post-Hoc Mitigation in Audio Deepfake Detection

语音伪造检测 | 6.6/10

 · 更新于 2026-09-06 · 约 20 分钟 · 9677 字 阅读 →
论文解读

Where Speech Enhancement Hurts Recognition: An Inference Time Polar Projection Diagnosis

语音识别 | 6.7/10

 · 更新于 2026-09-06 · 约 20 分钟 · 9645 字 阅读 →
论文解读

Which Languages Transfer Best to Warlpiri? A Similarity-Based Study for Low-Resource ASR

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6688 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-07-14

共分析 53 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 186 分钟 · 92760 字 阅读 →
论文解读

Beyond Time Shifts: Adapting Omni-LLM as a Reference-Free Evaluator for Generative Audio-Visual Models

音视频理解 | 6.0/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8333 字 阅读 →
论文解读

Clean2FX: Label-conditioned modeling for clean-to-effect guitar audio transformations

音频理解 | 7.3/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7034 字 阅读 →
论文解读

Dual-BEATs: Unlocking Zero-Shot Stereo Audio Perception in Audio Large Language Models via Dithering

多模态模型 | 7.1/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7057 字 阅读 →
论文解读

Event-Based Token Sequences for Audio-Conditioned Music-Game Level Modeling

音乐生成 | 7.2/10

 · 更新于 2026-09-06 · 约 18 分钟 · 8543 字 阅读 →
论文解读

FreyaTTS Technical Report

语音合成 | 7.7/10

 · 更新于 2026-09-06 · 约 18 分钟 · 8785 字 阅读 →
论文解读

Immersive Social Interaction with VR and LLM-Assisted Humanoids

语音交互 | 4.7/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5532 字 阅读 →
论文解读

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition

音视频语音识别 | 6.9/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7585 字 阅读 →
论文解读

Phone Segmentation and Recognition through Phonological Activation Mapping

语音识别 | 7.7/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8145 字 阅读 →