论文解读

PersonaGesture: Single-Reference Co-Speech Gesture Personalization for Unseen Speakers

音频生成 | 7.0/10

 · 更新于 2026-09-10 · 约 14 分钟 · 6606 字 阅读 →
论文解读

X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction

移动代理 | 6.5/10

 · 更新于 2026-09-10 · 约 12 分钟 · 5789 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-09

共分析 3 篇语音/AI 论文

 · 更新于 2026-09-10 · 约 10 分钟 · 4857 字 阅读 →
论文解读

Automated Clinical Report Generation for Remote Cognitive Remediation: Comparing Knowledge-Engineered Templates and LLMs in Low-Resource Settings

临床报告生成 | 7.5/10

 · 更新于 2026-09-10 · 约 25 分钟 · 12404 字 阅读 →
论文解读

Cross-Modal Navigation with Multi-Agent Reinforcement Learning

具身导航 | 7.5/10

 · 更新于 2026-09-10 · 约 14 分钟 · 6536 字 阅读 →
论文解读

Do Melody and Rhythm Coevolve?

音乐认知 | 7.5/10

 · 更新于 2026-09-10 · 约 22 分钟 · 10700 字 阅读 →
论文解读

Edge-specific signal propagation on mature chromophore-region 3D mechanism graphs for fluorescent protein quantum-yield prediction

蛋白质工程 | 7.5/10

 · 更新于 2026-09-10 · 约 13 分钟 · 6287 字 阅读 →
论文解读

Linear Semantic Segmentation for Low-Resource Spoken Dialects

语义分割 | 7.5/10

 · 更新于 2026-09-10 · 约 16 分钟 · 7887 字 阅读 →
论文解读

LiVeAction: a Lightweight, Versatile, and Asymmetric Neural Codec Design for Real-time Operation

多模态压缩 | 8.5/10

 · 更新于 2026-09-10 · 约 26 分钟 · 12573 字 阅读 →
论文解读

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM

语音大模型 | 8.0/10

 · 更新于 2026-09-10 · 约 29 分钟 · 14118 字 阅读 →
论文解读

Modality-Aware Contrastive and Uncertainty-Regularized Emotion Recognition

语音情感识别 | 8.0/10

 · 更新于 2026-09-10 · 约 16 分钟 · 7828 字 阅读 →
论文解读

More Than Can Be Said: A Benchmark and Framework for Pre-Question Scientific Ideation

基准测试 | 6.5/10

 · 更新于 2026-09-10 · 约 9 分钟 · 4119 字 阅读 →
论文解读

MultiLinguahah : A New Unsupervised Multilingual Acoustic Laughter Segmentation Method

音频事件检测 | 8.5/10

 · 更新于 2026-09-10 · 约 13 分钟 · 6396 字 阅读 →
论文解读

NDF+: Joint Neural Directional Filtering and Diffuse Sound Extraction

空间音频 | 6.5/10

 · 更新于 2026-09-10 · 约 11 分钟 · 5284 字 阅读 →
论文解读

Optimal Transport Audio Distance with Learned Riemannian Ground Metrics

音频质量评估 | 7.0/10

 · 更新于 2026-09-10 · 约 18 分钟 · 8844 字 阅读 →
论文解读

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization

音频编码 | 7.0/10

 · 更新于 2026-09-10 · 约 14 分钟 · 6791 字 阅读 →
论文解读

PersonaKit (PK): A Plug-and-Play Platform for User Testing Diverse Roles in Full-Duplex Dialogue

全双工对话系统评估 | 6.0/10

 · 更新于 2026-09-10 · 约 24 分钟 · 11903 字 阅读 →
论文解读

PianoCoRe: Combined and Refined Piano MIDI Dataset

数据集 | 7.5/10

 · 更新于 2026-09-10 · 约 16 分钟 · 7756 字 阅读 →
论文解读

Predictive-Generative Drift Decomposition for Speech Enhancement and Separation

语音增强 | 8.5/10

 · 更新于 2026-09-10 · 约 18 分钟 · 8741 字 阅读 →
论文解读

Preliminary Insights in Chronos Frequency Data Understanding and Reconstruction

模型评估 | 6.0/10

 · 更新于 2026-09-10 · 约 15 分钟 · 7215 字 阅读 →