论文解读

I Understand How You Feel: Enhancing Deeper Emotional Support Through Multilingual Emotional Validation in Dialogue System

语音识别 | 5.8/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7227 字 阅读 →
论文解读

Massive Open-Vocabulary Keyword Spotting

语音识别 | 9.8/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5180 字 阅读 →
论文解读

Pretrained self-supervised speech models can recognize unseen consonants

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4929 字 阅读 →
论文解读

SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations

语音合成 | 7.9/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4632 字 阅读 →
论文解读

Towards Data-free and Training-free Compression for Speech Foundation Models Using Parameter Clustering

语音识别 | 6.4/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5326 字 阅读 →
论文解读

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6593 字 阅读 →
论文解读

AuRA: Internalizing Audio Understanding into LLMs as LoRA

语音问答 | 7.5/10

 · 更新于 2026-09-25 · 约 7 分钟 · 3100 字 阅读 →
论文解读

Entropy-Aware Domain-Routed Mixture-of-Experts Speech-LLM Framework: A Case Study of Multi-Domain Child-Adult ASR

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6225 字 阅读 →
论文解读

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6823 字 阅读 →
论文解读

GC-LoRA: Gated Convolutional LoRA for Parameter-Efficient Acoustic Adaptation

语音识别 | 7.6/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4079 字 阅读 →
论文解读

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling

语音识别 | 8.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5701 字 阅读 →
论文解读

OmniCap-IF: Benchmarking and Improving Instruction Following Abilities for Omni-Video Captioning

语音生成 | 9.1/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6024 字 阅读 →
论文解读

Phoneme-First Prediction for LLM-Based Speech Recognition

语音识别 | 6.9/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6037 字 阅读 →
论文解读

Speaker Group Encoding in Self-supervised Speech Recognition Models

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5916 字 阅读 →
论文解读

Speech Encoder Fusion for LLM-based Automatic Speech Recognition

语音识别 | 7.2/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5281 字 阅读 →
论文解读

Speech Meets ELF: Audio Conditional Continuous-Target Diffusion for Speech Recognition and Translation

语音识别 | 8.3/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6269 字 阅读 →
论文解读

Towards Deep Contextual Reasoning from Broad Descriptions for ASR with Speech-LLM via Metadata-Driven Reasoning Chains

语音识别 | 6.2/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4631 字 阅读 →
论文解读

TRADE: Transducer-Augmented Decoder for Speech LLM

语音识别 | 7.4/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5157 字 阅读 →
论文解读

ViP-VL: Vietnamese Self-supervised Speech Pretraining Model with Vector-Quantization Learning

语音识别 | 9.7/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5467 字 阅读 →
论文解读

A study on the impact of region specific data on the performance of Indic ASR

语音识别 | 7.2/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6491 字 阅读 →