论文解读

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling

语音识别 | 8.6/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5701 字 阅读 →
论文解读

OmniCap-IF: Benchmarking and Improving Instruction Following Abilities for Omni-Video Captioning

语音生成 | 9.1/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6024 字 阅读 →
论文解读

Phoneme-First Prediction for LLM-Based Speech Recognition

语音识别 | 6.9/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6037 字 阅读 →
论文解读

Speaker Group Encoding in Self-supervised Speech Recognition Models

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5916 字 阅读 →
论文解读

Speech Encoder Fusion for LLM-based Automatic Speech Recognition

语音识别 | 7.2/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5281 字 阅读 →
论文解读

Speech Meets ELF: Audio Conditional Continuous-Target Diffusion for Speech Recognition and Translation

语音识别 | 8.3/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6269 字 阅读 →
论文解读

Towards Deep Contextual Reasoning from Broad Descriptions for ASR with Speech-LLM via Metadata-Driven Reasoning Chains

语音识别 | 6.2/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4631 字 阅读 →
论文解读

TRADE: Transducer-Augmented Decoder for Speech LLM

语音识别 | 7.4/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5157 字 阅读 →
论文解读

ViP-VL: Vietnamese Self-supervised Speech Pretraining Model with Vector-Quantization Learning

语音识别 | 9.7/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5467 字 阅读 →
论文解读

A study on the impact of region specific data on the performance of Indic ASR

语音识别 | 7.2/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6491 字 阅读 →
论文解读

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs

语音识别 | 8.8/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6046 字 阅读 →
论文解读

Bridging Traditional Explainability Methods and Multimodal Multilingual Models: An XAI-Based Analysis

语音识别 | 5.4/10

 · 更新于 2026-09-06 · 约 18 分钟 · 8872 字 阅读 →
论文解读

Cross-Modal Masking for Robust Silent Speech Synthesis Using sEMG and Lipreading

语音合成 | 7.5/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7716 字 阅读 →
论文解读

Factors affecting ASR performance: A study using state of the art ASR models in Indic Languages

语音识别 | 6.9/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5995 字 阅读 →
论文解读

Is Text All You Need? Text as a Universal Information Bottleneck for Speech LLMs

语音识别 | 7.6/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5311 字 阅读 →
论文解读

Overcoming Decoder Inconsistencies in Whisper for Dravidian and Low-Resource Languages

语音识别 | 6.2/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6234 字 阅读 →
论文解读

Parameter-Efficient Continual Learning for Automatic Speech Recognition

语音识别 | 8.1/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5336 字 阅读 →
论文解读

Subtitle-Aligned Fine-Tuning of Whisper for Swiss German ASR: Benchmark Contamination, Convention Mismatch, and an Honest Baseline at 25.6% WER (13.8% cWER)

语音识别 | 8.2/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6640 字 阅读 →
论文解读

Assessing True Generalisability of Audio-Visual Speech Recognisers

语音识别 | 9.5/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6841 字 阅读 →
论文解读

Contrastive Training with LLM-generated Near-Misses for Robust Code-Switching Speech Recognition

语音识别 | 7.1/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5503 字 阅读 →