论文解读

A Lightweight Dual-Factor Acoustic Authentication System via Cascaded GMM-DTW Architecture for Edge Computing

说话人验证 | 6/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6198 字 阅读 →
论文解读

Automated Pronunciation Evaluation for Korean Toddler Speech using Speech Diarization and Self-Supervised Learning

说话人日志 | 6/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6726 字 阅读 →
论文解读

Entropy-Aware Domain-Routed Mixture-of-Experts Speech-LLM Framework: A Case Study of Multi-Domain Child-Adult ASR

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6225 字 阅读 →
论文解读

GlobeAudio: A Multilingual Multicultural Benchmark for Naturalistic Evaluation of Large Audio-Language Models

语音识别 | 7.9/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5486 字 阅读 →
论文解读

Phoneme-First Prediction for LLM-Based Speech Recognition

语音识别 | 6.9/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6037 字 阅读 →
论文解读

Recovering the Zipfian Distribution in Unsupervised Term Discovery

自监督学习 | 8.7/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5173 字 阅读 →
论文解读

Speech Encoder Fusion for LLM-based Automatic Speech Recognition

语音识别 | 7.2/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5281 字 阅读 →
论文解读

SSL-GMMVC: Interpretable Voice Conversion via Locally Linear GMM Transforms in Self-Supervised Representation Space

语音转换 | 6.8/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5852 字 阅读 →
论文解读

Towards Robust Arabic Speech Emotion Recognition with Deep Learning

语音情感识别 | 6.4/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5530 字 阅读 →
论文解读

ViP-VL: Vietnamese Self-supervised Speech Pretraining Model with Vector-Quantization Learning

语音识别 | 9.7/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5467 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-10

共分析 45 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 126 分钟 · 62690 字 阅读 →
论文解读

A Comparative Study of Pre-trained Speech Encoders and Training Objectives for Large-Scale Indic Spoken Language Identification

自监督学习 | 8.9/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4667 字 阅读 →
论文解读

A study on the impact of region specific data on the performance of Indic ASR

语音识别 | 7.2/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6491 字 阅读 →
论文解读

Factors affecting ASR performance: A study using state of the art ASR models in Indic Languages

语音识别 | 6.9/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5995 字 阅读 →
论文解读

NüshuVoice: Reviving the Voice of Endangered Nüshu with Pitch-Aware Text-to-Speech

语音合成 | 7/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5298 字 阅读 →
论文解读

OpenBibleTTS: Large-Scale Speech Resources and TTS Models for Low-Resource Languages

语音合成 | 8/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6802 字 阅读 →
论文解读

Overcoming Decoder Inconsistencies in Whisper for Dravidian and Low-Resource Languages

语音识别 | 6.2/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6234 字 阅读 →
论文解读

Parameter-Efficient Continual Learning for Automatic Speech Recognition

语音识别 | 8.1/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5336 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-09

共分析 48 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 141 分钟 · 70154 字 阅读 →
论文解读

dots.tts Technical Report

语音合成 | 9/10

 · 更新于 2026-09-06 · 约 7 分钟 · 3032 字 阅读 →