论文解读

Evaluating Music Context Preservation: A Multi-facet Framework for Music Editing Systems

音乐理解 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5074 字 阅读 →
论文解读

Finetuning Strategies for Querying Sounds by Vocal Imitation

音频检索 | 7.3/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5013 字 阅读 →
论文解读

FM Synthesizer Audio-Parameter Shared Embeddings

音频生成 | 7.3/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4702 字 阅读 →
论文解读

Generalized Audio-Driven Synthesis of Precise Drummer Motion

音视频生成 | 7.8/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4876 字 阅读 →
论文解读

Geometric Iterative Retrieval for Neural Audio Codec Resynthesis

音频编码 | 7.6/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5287 字 阅读 →
论文解读

Large Language Models in Mental Health: A Systematic Review of Applications, Innovations, and Ethical Challenges

多模态模型 | 6.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4562 字 阅读 →
论文解读

Low-Power, Neuromorphic, Acoustic Anomaly Detection for Persistent Machine Monitoring

音频分类 | 7.3/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5052 字 阅读 →
论文解读

Mitigating Spectral Bias in Neural Operators for Underwater Transmission Loss Prediction

音频理解 | 7.3/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4170 字 阅读 →
论文解读

Multimodal Rapport Estimation in Real-World HRI

多模态模型 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4533 字 阅读 →
论文解读

Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities

音频理解 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4569 字 阅读 →
论文解读

Pedagogical AI in Mental Health: A Tri-Stream Fine-Tuned LLM Framework for Automated Clinical Supervision and Risk Triage

多模态模型 | 6.7/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4594 字 阅读 →
论文解读

Sounds Uncertain: Exploring the Affective Aspects of Sonification for Uncertainty Visualization

音频生成 | 6.7/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4366 字 阅读 →
论文解读

StocksTalk: A Voice-Enabled Conversational Agent for Structured Query Generation over Web Data

语音交互 | 6.7/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4396 字 阅读 →
论文解读

Understanding Multilingual Medical ASR Adaptation Through Layer-Wise Analysis

语音识别 | 7.2/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5145 字 阅读 →
论文解读

X2Streaming-TTS: Causal Token-Level Text-to-Speech from Streaming Text with Speech-State Inheritance

语音合成 | 7.9/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4444 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-08-20

共分析 20 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 47 分钟 · 23525 字 阅读 →
论文解读

A Multiplication-Free Feature Extractor for Signal Classification: Keyword Spotting Case Study

语音唤醒 | 6.9/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7078 字 阅读 →
论文解读

Automatic Transcription of Microtonal Free-Rhythm Vocal Music: A Case Study in Iranian Classical Music

音乐转录 | 7.5/10

 · 更新于 2026-09-06 · 约 7 分钟 · 3026 字 阅读 →
论文解读

Closing the Affective Loop: Multimodal Speaker-Listener Emotion-Dynamics-Aware Empathetic Social Robots

语音交互 | 6.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3605 字 阅读 →
论文解读

DNN-Based Frequency-Dependent Estimation of Speech, Music, and Noise Power in Acoustic Mixtures for Hearing-Aid Scene Analysis

音频分离 | 6.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4890 字 阅读 →