论文解读

Whisper-Aware LLM: Self-Supervised Uncertainty Learning for Robust Whispered Speech Recognition

语音识别 | 6.1/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6535 字 阅读 →
论文解读

AI-Guided Learning: Research on Knowledge and Skill Acquisition Support Methods Using Deep Learning Audio-Video Processing Techniques

音视频理解 | 6.3/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6962 字 阅读 →
论文解读

PACE: A Playback-Aligned Context Engine for LLM-Based Full-Duplex Voice Dialogue

语音交互 | 8.0/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6734 字 阅读 →
论文解读

SraVaani 1.0: Scaling Inclusive Speech Recognition for Indic Languages

语音识别 | 7.9/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8512 字 阅读 →
论文解读

LSEAD: A Privacy-Preserving LLM-Based Speech Analysis Framework for Early Alzheimer's Disease Screening

语音识别 | 6.6/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7355 字 阅读 →
论文解读

Decolonizing Linguistic Policies in Automated Speech Recognition: A Framework for Cross-Culturally Competent Speech AI

语音识别 | 7.8/10

 · 更新于 2026-09-25 · 约 7 分钟 · 3465 字 阅读 →
论文解读

ECHO: A Locally-Deployable Agentic Health Assistant with Temporal Memory, Safety Guardrails, and Speech Assessment

语音属性识别 | 5.5/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7091 字 阅读 →
论文解读

FormBharo: Designing and Evaluating a Voice Agent for Conversational Form Filling in Rural India

语音交互 | 7.0/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6937 字 阅读 →
论文解读

How to Recognize New Words: A Comparison Between Context Biasing Methods and Speech LLMs

语音识别 | 6.4/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7408 字 阅读 →
论文解读

Language-Specialized Multi-Teacher On-Policy Distillation for Multilingual LLM-Based ASR

语音识别 | 6.9/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6440 字 阅读 →
论文解读

Latent Softmax for Data-Efficient Phoneme-Based Multilingual ASR Across Tonal and Non-Tonal Languages

语音识别 | 6.3/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5826 字 阅读 →
论文解读

Analyzing Speech Condition Effects in Dysarthric ASR: A Layer-wise Probing Study

语音识别 | 5.0/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6341 字 阅读 →
论文解读

Embodied Empathy: A Multimodal AR and LLM-Powered System for Self-Attachment Psychotherapy with Self-Initiated Humour

音频交互 | 5.4/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6236 字 阅读 →
论文解读

Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval

音频检索 | 7.9/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8227 字 阅读 →
论文解读

Normal-Anchored First-Order Model-Agnostic Meta-Learning based Whisper Fine-Tuning for Enhancing Fairness of Cleft Lip and Palate Speech Recognition

语音识别 | 5.9/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6336 字 阅读 →
论文解读

SoniSpeech: A Large-Scale Open-Vocabulary Tri-Modal Dataset for Wearable Silent Speech Interfaces

语音识别 | 8.3/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5919 字 阅读 →
论文解读

Leveraging Beam Search Information for Confidence Estimation in E2E ASR

语音识别 | 6.3/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7632 字 阅读 →
论文解读

ParaASR: Multi-Token Prediction for Fast and Long-Context LLM-Based Speech Recognition

语音识别 | 6.7/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5685 字 阅读 →
论文解读

AgenticASR: Refining Speech Recognition in Real-World Scenarios via an Agentic Approach

语音识别 | 6.8/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6977 字 阅读 →
论文解读

Does EEG Foundation Models Transfer to Speech? A Benchmark on Overt and Imagined Speech Decoding

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5912 字 阅读 →