论文解读

MusiChat: Vibe Composing for Music Creation

音乐生成 | 5.8/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6218 字 阅读 →
论文解读

MyMentorLLM: A psychotherapy GenAI environment with multimodal voice/text patients, trainees and experts for deliberate practice

语音交互 | 5.9/10

 · 更新于 2026-09-25 · 约 20 分钟 · 9675 字 阅读 →
论文解读

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models

音视频问答 | 8.3/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8316 字 阅读 →
论文解读

Parallel Decoding Distillation for Fast Image and Video Generation

音视频生成 | 6.2/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9236 字 阅读 →
论文解读

Spacing Out: On the Reliability of Binaural Music Source Separation Metrics

音乐源分离 | 6.1/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6032 字 阅读 →
论文解读

SpeechLLM Meets Federated Learning for End-to-End ASR: English and Italian Case Studies

语音识别 | 7.1/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5520 字 阅读 →
论文解读

Text-Prompted CLAP: Learning Query-Conditioned Audio Representations via Contrastive Learning

音频理解 | 6.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5915 字 阅读 →
论文解读

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models

声源定位 | 7.2/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8573 字 阅读 →
论文解读

VAD to the Bone: Ultra-Tiny Speech Activity Detection for Edge Deployment

语音活动检测 | 7.4/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6296 字 阅读 →
论文解读

Automatic Audio Equalization with Semantic Embeddings

语音增强 | 5.8/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7531 字 阅读 →
论文解读

Do Visual Features Improve Other-Initiated Repair Detection? A Dyadic Multimodal Approach

音视频交互 | 5.9/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7206 字 阅读 →
论文解读

Earnings25: A Comprehensive 500-Hour Speech Benchmark for Finance

语音识别 | 7.1/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5808 字 阅读 →
论文解读

Explainable AI through the Lens of Material Agency: Enabling Musical Interface Design with Neural Audio Models

音乐生成 | 6.7/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6159 字 阅读 →
论文解读

Expose Your Disguise: Recovering Source Speaker Identity From Voice Conversion

对比学习 | 7.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6629 字 阅读 →
论文解读

Improving Zero-Shot Phonetic Classification through Language-Agnostic Articulatory Features

语音属性识别 | 5.3/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7242 字 阅读 →
论文解读

Infinite Canons: Maximally Self-Similar Melodic Lines and Canons with Infinite Solutions

音乐生成 | 6.4/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6656 字 阅读 →
论文解读

JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents

音视频生成 | 6.2/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6160 字 阅读 →
论文解读

Leveraging Gradient Reversal Loss and Multitask Learning for Datasets-Aware Audio Deepfake Detection

语音伪造检测 | 7.3/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6744 字 阅读 →
论文解读

Looking for Affect in Spontaneous Finnish Speech through Linguistic Interpretability

语音情感识别 | 5.3/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5152 字 阅读 →
论文解读

Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers

语音合成 | 6.8/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7089 字 阅读 →