论文解读

Beyond the Mouth: Upper-Face Affective Cues in Audiovisual Sentence Recognition under Acoustic Uncertainty

语音识别 | 5.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4920 字 阅读 →
论文解读

Echo: A Joint-Embedding Predictive Architecture for Speaker Diarization and Speech Recognition in a Shared Latent Space

语音识别 | 7/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7107 字 阅读 →
论文解读

MOSS-Audio Technical Report

语音识别 | 9.2/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5940 字 阅读 →
论文解读

MURMUR: An Efficient Inference System for Long-Form ASR

语音识别 | 8.3/10

 · 更新于 2026-09-25 · 约 4 分钟 · 1773 字 阅读 →
论文解读

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects

语音识别 | 8.8/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5668 字 阅读 →
论文解读

SN-WER: Script-Normalized WER for Multi-Script Indic ASR Evaluation

语音识别 | 5.3/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6052 字 阅读 →
论文解读

Spiking and Event-driven Neuromorphic Mamba Models for Efficient Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6692 字 阅读 →
论文解读

Sympatheia: Emotionally Adaptive Voice Assistant with Continuous Affect Conditioning

语音合成 | 9.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5542 字 阅读 →
论文解读

WAXAL-NET: Finetuned Edge ASR Across 19 African Languages

语音识别 | 8/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5632 字 阅读 →
论文解读

A Unified and Reproducible Experimentation Framework for Speech Understanding

语音识别 | 5.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6336 字 阅读 →
论文解读

MindVoice: Reconstructing Intelligible Speech from Non-invasive Neural Signals with Pretrained Priors

语音合成 | 8.5/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8301 字 阅读 →
论文解读

Scaling Conversational Hungarian ASR: The BEA-Dialogue+ Corpus

语音识别 | 7.2/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5792 字 阅读 →
论文解读

UniAudio-Token: Empowering Semantic Speech Tokenizers with General Audio Perception

语音合成 | 10/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7690 字 阅读 →
论文解读

Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs

语音识别 | 7.2/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5495 字 阅读 →
论文解读

Raon-Speech Technical Report

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5464 字 阅读 →
论文解读

Archon: A Unified Multimodal Model for Holistic Digital Human Generation

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5475 字 阅读 →
论文解读

ChildVox: A Speech, Audio, and Large Audio-Language Model Benchmark in Understanding and Characterizing Sound across Childhood

语音识别 | 8/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5167 字 阅读 →
论文解读

Data-Efficient On-Policy Distillation for Automatic Speech Recognition

语音识别 | 5.1/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4830 字 阅读 →
论文解读

Dial HEALTHDIAL for Advice: A Multilingual and Multi-Parallel Spoken Dialogue Dataset for Knowledge-Grounded Information Seeking

语音合成 | 8.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5895 字 阅读 →
论文解读

HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding

语音合成 | 8.6/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6179 字 阅读 →