论文解读

LightAVSeg: Lightweight Audio-Visual Segmentation

模型压缩 | 6.3/10

 · 更新于 2026-09-24 · 约 18 分钟 · 8803 字 阅读 →
论文解读

REST: Diffusion-based Real-time End-to-end Streaming Talking Head Generation via ID-Context Caching and Asynchronous Streaming Distillation

音视频生成 | 7.3/10

 · 更新于 2026-09-24 · 约 26 分钟 · 12601 字 阅读 →
论文解读

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning

知识蒸馏 | 7/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6068 字 阅读 →
论文解读

A Multi-Branch Hierarchy-Aware Framework for Heterogeneous Audio Classification

音频分类 | 4.9/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7182 字 阅读 →
论文解读

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation

知识蒸馏 | 10/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5503 字 阅读 →
论文解读

Thinking While Speaking: Inference-Time Knowledge Transfer for Responsive and Intelligent Conversational Voice Agents

知识蒸馏 | 6.7/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5061 字 阅读 →
论文解读

End-to-End Voice Intent Recognition for Spontaneous Human-Drone Interaction with Naive Users

端到端 | 7/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6231 字 阅读 →
论文解读

Speaker Identity in Non-Verbal Vocalizations: Conditional Distillation and Mixture of Experts Approach

说话人验证 | 9.1/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6090 字 阅读 →
论文解读

Streaming T5-based Text-to-Speech Synthesis with Limited Lookahead

语音合成 | 6.7/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5591 字 阅读 →
论文解读

Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs

语音合成 | 7.4/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7228 字 阅读 →
论文解读

Teacher-Student Structure for Domain Adaptation in Ensemble Audio-Visual Video Deepfake Detection

多模态模型 | 7.4/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5649 字 阅读 →
论文解读

Decoding Insect Song: A Multitask Semisupervised Orthoptera Bioacoustic Classifier

音频分类 | 8.7/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6708 字 阅读 →
论文解读

Missing-Token Prompted Reliability-Aware Fusion for Robust Polyglot Speaker Identification

说话人识别 | 8.6/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5723 字 阅读 →
论文解读

Fast Speech Foundation Model Distillation Using Interleaved Stacking

知识蒸馏 | 6.6/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4840 字 阅读 →
论文解读

AuRA: Internalizing Audio Understanding into LLMs as LoRA

语音问答 | 7.5/10

 · 更新于 2026-09-24 · 约 7 分钟 · 3100 字 阅读 →
论文解读

Cross-Modal Knowledge Distillation without Paired Data: Theoretical Foundation and Algorithm

语音识别 | 7.5/10

 · 更新于 2026-09-24 · 约 21 分钟 · 10471 字 阅读 →
论文解读

USAD 2.0: Scaling Representation Distillation for Universal Audio Understanding

音频编码 | 9/10

 · 更新于 2026-09-24 · 约 12 分钟 · 6009 字 阅读 →
论文解读

Logit Distillation on Manifolds: Mapping by Learning

语音识别 | 6.5/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6262 字 阅读 →
论文解读

Raon-Speech Technical Report

语音识别 | 6.5/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5464 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-30

共分析 6 篇语音/AI 论文

 · 更新于 2026-09-24 · 约 17 分钟 · 8127 字 阅读 →