论文解读

GraphIDyOM: A graph-native Python reimplementation of IDyOM for musical expectation modelling

音乐理解 | 7.8/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5997 字 阅读 →
论文解读

Joint Text-Audio Alignment for EEG-to-Text Decoding in Chinese Speech Production and Perception

语音交互 | 7.0/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7535 字 阅读 →
论文解读

Multi-Phonation Graph Learning with Self-Supervised Speech Embeddings for ALS Detection and Progression Prediction

语音属性识别 | 5.5/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6989 字 阅读 →
论文解读

MusiChat: Vibe Composing for Music Creation

音乐生成 | 5.8/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6218 字 阅读 →
论文解读

MyMentorLLM: A psychotherapy GenAI environment with multimodal voice/text patients, trainees and experts for deliberate practice

语音交互 | 5.9/10

 · 更新于 2026-09-06 · 约 20 分钟 · 9675 字 阅读 →
论文解读

OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs

音视频问答 | 7.0/10

 · 更新于 2026-09-06 · 约 19 分钟 · 9455 字 阅读 →
论文解读

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models

音视频问答 | 8.3/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8316 字 阅读 →
论文解读

Parallel Decoding Distillation for Fast Image and Video Generation

音视频生成 | 6.2/10

 · 更新于 2026-09-06 · 约 19 分钟 · 9236 字 阅读 →
论文解读

Self-Supervised Audio Representation Learning for Pediatric Asthma Detection in Emergency Care Using Digital Stethoscope Recordings

音频分类 | 5.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5443 字 阅读 →
论文解读

Spacing Out: On the Reliability of Binaural Music Source Separation Metrics

音乐源分离 | 6.1/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6032 字 阅读 →
论文解读

SpeechLLM Meets Federated Learning for End-to-End ASR: English and Italian Case Studies

语音识别 | 7.1/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5520 字 阅读 →
论文解读

Text-Prompted CLAP: Learning Query-Conditioned Audio Representations via Contrastive Learning

音频理解 | 6.6/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5915 字 阅读 →
论文解读

Towards Operational Conversational Intelligence: A Speech Intelligence Framework

说话人日志 | 6.4/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6520 字 阅读 →
论文解读

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models

声源定位 | 7.2/10

 · 更新于 2026-09-06 · 约 18 分钟 · 8573 字 阅读 →
论文解读

VAD to the Bone: Ultra-Tiny Speech Activity Detection for Edge Deployment

语音活动检测 | 7.4/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6296 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-07-29

共分析 27 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 83 分钟 · 41435 字 阅读 →
论文解读

Automatic Audio Equalization with Semantic Embeddings

语音增强 | 5.8/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7531 字 阅读 →
论文解读

Disentangling Acoustic Cues in Alzheimer's Pathology and Perception: The Roles of Language and Gender

语音属性识别 | 5.4/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7657 字 阅读 →
论文解读

Do Visual Features Improve Other-Initiated Repair Detection? A Dyadic Multimodal Approach

音视频交互 | 5.9/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7206 字 阅读 →
论文解读

Earnings25: A Comprehensive 500-Hour Speech Benchmark for Finance

语音识别 | 7.1/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5808 字 阅读 →