论文解读

C2GA: A Class-Controllable Generative Augmentation Framework for Respiratory Sound Classification

音频分类 | 7.3/10

 · 更新于 2026-09-25 · 约 6 分钟 · 2877 字 阅读 →
论文解读

CoughSense: Five-Class Respiratory Disease Classification via Whisper Encoder Fine-Tuning and Dual-Encoder Cross-Attention Fusion with Balanced Contrastive Learning

数据增强 | 9.1/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7875 字 阅读 →
论文解读

Efficient ASR Training with Conversations that Never Happened

语音识别 | 8/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6552 字 阅读 →
论文解读

SegTune: Structured and Fine-Grained Control for Song Generation

音乐生成 | 8.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6788 字 阅读 →
论文解读

SiamCTC: Learning Speech Representations through Monotonic Temporal Alignment

语音识别 | 7/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6619 字 阅读 →
论文解读

SketchSong: Hierarchical Song Generation with Sketch Planning and Fine-Grained Multi-Track Modeling

音乐生成 | 8.6/10

 · 更新于 2026-09-25 · 约 26 分钟 · 12623 字 阅读 →
论文解读

Advancing Electrolaryngeal Speech Enhancement Through Speech-Text Representation Learning

语音增强 | 7.1/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8092 字 阅读 →
论文解读

AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling

多模态模型 | 7/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7267 字 阅读 →
论文解读

Echo: A Joint-Embedding Predictive Architecture for Speaker Diarization and Speech Recognition in a Shared Latent Space

语音识别 | 7/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7107 字 阅读 →
论文解读

UniVocal: Unified Speech-Singing Code-Switching Synthesis

语音合成 | 8.9/10

 · 更新于 2026-09-25 · 约 5 分钟 · 2097 字 阅读 →
论文解读

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment

语音合成 | 9.3/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6263 字 阅读 →
论文解读

Improving acoustic drone detection generalization through pretraining and data augmentation

音频事件检测 | 7.7/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5246 字 阅读 →
论文解读

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue

语音合成 | 8.9/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8091 字 阅读 →
论文解读

Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs

语音识别 | 7.2/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5495 字 阅读 →
论文解读

EchoDistill:Alignment Noisy-to-Clean Self-Distillation for Robust Audio LLMs

强化学习 | 9.1/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6317 字 阅读 →
论文解读

Raon-Speech Technical Report

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5464 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-30

共分析 6 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 17 分钟 · 8127 字 阅读 →
论文解读

Archon: A Unified Multimodal Model for Holistic Digital Human Generation

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5475 字 阅读 →
论文解读

Audio Jailbreaks in Large Audio-Language Models: Taxonomy, Attack-Defense Analysis, and Cost-Aware Evaluation

多模态模型 | 8.9/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6155 字 阅读 →
论文解读

Dial HEALTHDIAL for Advice: A Multilingual and Multi-Parallel Spoken Dialogue Dataset for Knowledge-Grounded Information Seeking

语音合成 | 8.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5895 字 阅读 →