论文解读

MindVoice: Reconstructing Intelligible Speech from Non-invasive Neural Signals with Pretrained Priors

语音合成 | 8.5/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8301 字 阅读 →
论文解读

Scaling Conversational Hungarian ASR: The BEA-Dialogue+ Corpus

语音识别 | 7.2/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5792 字 阅读 →
论文解读

UniAudio-Token: Empowering Semantic Speech Tokenizers with General Audio Perception

语音合成 | 10/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7690 字 阅读 →
论文解读

Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs

语音识别 | 7.2/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5495 字 阅读 →
论文解读

Raon-Speech Technical Report

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5464 字 阅读 →
论文解读

Archon: A Unified Multimodal Model for Holistic Digital Human Generation

语音合成 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5475 字 阅读 →
论文解读

ChildVox: A Speech, Audio, and Large Audio-Language Model Benchmark in Understanding and Characterizing Sound across Childhood

语音识别 | 8/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5167 字 阅读 →
论文解读

Data-Efficient On-Policy Distillation for Automatic Speech Recognition

语音识别 | 5.1/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4830 字 阅读 →
论文解读

Dial HEALTHDIAL for Advice: A Multilingual and Multi-Parallel Spoken Dialogue Dataset for Knowledge-Grounded Information Seeking

语音合成 | 8.6/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5895 字 阅读 →
论文解读

HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding

语音合成 | 8.6/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6179 字 阅读 →
论文解读

MELD: Mel-Spectrogram-Based Speech Language Modeling with Discrete Latent Variables

语音合成 | 7.3/10

 · 更新于 2026-09-06 · 约 3 分钟 · 1272 字 阅读 →
论文解读

OmniInteract: Benchmarking Real-World Streaming Interaction for Real-Time Omnimodal Assistants

语音识别 | 7.8/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6454 字 阅读 →
论文解读

The WER Trap: Shattering the Illusion of Unified Tokens in Speech Language Models

语音识别 | 7/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6823 字 阅读 →
论文解读

VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents

语音合成 | 7.9/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7179 字 阅读 →
论文解读

Benchmarking AI for low-resource contexts: Thinking beyond leaderboards

语音识别 | 5.1/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6151 字 阅读 →
论文解读

Breaking the Script Barrier: Enabling Automatic Alignment for PoS-based ASR Error Analysis in Non-Latin Scripts

语音识别 | 6.0/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6823 字 阅读 →
论文解读

Building Community-Centred NLP Resources for Puno Quechua

语音识别 | 7.2/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5330 字 阅读 →
论文解读

Do Audio LLMs Listen or Read? Analyzing and Mitigating Paralinguistic Failures with VoxParadox

语音情感识别 | 6.8/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5436 字 阅读 →
论文解读

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini

语音识别 | 7.9/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5387 字 阅读 →
论文解读

When Helpful Context Leaks: Privacy Risks in Domain-Adapted ASR

语音识别 | 10/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6313 字 阅读 →