论文解读

Automated Clinical Report Generation for Remote Cognitive Remediation: Comparing Knowledge-Engineered Templates and LLMs in Low-Resource Settings

临床报告生成 | 7.5/10

 · 更新于 2026-09-25 · 约 25 分钟 · 12404 字 阅读 →
论文解读

Linear Semantic Segmentation for Low-Resource Spoken Dialects

语义分割 | 7.5/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7887 字 阅读 →
论文解读

More Than Can Be Said: A Benchmark and Framework for Pre-Question Scientific Ideation

基准测试 | 6.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4119 字 阅读 →
论文解读

Pro-KLShampoo: Projected KL-Shampoo with Whitening Recovered by Orthogonalization

大语言模型 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4572 字 阅读 →
论文解读

Benchmarking LLMs on the Massive Sound Embedding Benchmark (MSEB)

音频分类 | 7.0/10

 · 更新于 2026-09-25 · 约 5 分钟 · 2190 字 阅读 →
论文解读

JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions

音频质量评估 | 8.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6589 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-07

共分析 22 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 79 分钟 · 39450 字 阅读 →
论文解读

DECKER: Domain-invariant Embedding for Cross-Keyboard Extraction and Recognition

音频安全 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5862 字 阅读 →
论文解读

Flexi-LoRA with Input-Adaptive Ranks: Efficient Finetuning for Speech and Reasoning Tasks

大语言模型 | 8.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5782 字 阅读 →
论文解读

MindMelody: A Closed-Loop EEG-Driven System for Personalized Music Intervention

音乐生成 | 7.0/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5860 字 阅读 →
论文解读

Virtual Speech Therapist: A Clinician-in-the-Loop AI Speech Therapy Agent for Personalized and Supervised Therapy

语音治疗系统 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4453 字 阅读 →
论文解读

AUHead: Realistic Emotional Talking Head Generation via Action Units Control

生成模型 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4862 字 阅读 →
论文解读

Closing the Gap Between Text and Speech Understanding in LLMs

语音大模型 | 8.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4518 字 阅读 →
论文解读

Compose and Fuse: Revisiting the Foundational Bottlenecks in Multimodal Reasoning

多模态推理 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4438 字 阅读 →
论文解读

Confident and Adaptive Generative Speech Recognition via Risk Control

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4401 字 阅读 →
论文解读

End-to-end Listen, Look, Speak and Act

语音对话系统 | 8.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5323 字 阅读 →
论文解读

From Text to Talk: Audio-Language Model Needs Non-Autoregressive Joint Training

语音对话系统 | 8.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5679 字 阅读 →
论文解读

Instilling an Active Mind in Avatars via Cognitive Simulation

音视频 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4426 字 阅读 →
论文解读

JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation

音视频联合推理 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4251 字 阅读 →
论文解读

LLM2Fx-Tools: Tool Calling for Music Post-Production

音乐信息检索 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4633 字 阅读 →