论文解读

One Voice, Many Tongues: Cross-Lingual Voice Cloning for Scientific Speech

语音克隆 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4425 字 阅读 →
论文解读

Step-Audio-R1.5 Technical Report

语音对话系统 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4318 字 阅读 →
论文解读

Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5056 字 阅读 →
论文解读

Advancing LLM-Based Multi-Channel Multi-Speaker Speech Recognition with Global Cross-Channel Attention and Sentence-Ordered First-In First-Out Serialized Output Training

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5058 字 阅读 →
论文解读

Advancing Speech Understanding in Speech-Aware Language Models with GRPO

语音问答 | 7.0/10

 · 更新于 2026-09-25 · 约 7 分钟 · 3396 字 阅读 →
论文解读

Adversarial Fine-Tuning on Speech Foundation Model with Vulnerable Attention Consistency Regularization for Robust Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4509 字 阅读 →
论文解读

Aligning Generative Speech Enhancement with Perceptual Feedback

语音增强 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5533 字 阅读 →
论文解读

Attention-Weighted Centered Kernel Alignment for Knowledge Distillation in Large Audio-Language Models Applied To Speech Emotion Recognition

语音情感识别 | 8.0/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5893 字 阅读 →
论文解读

Attention2Probability: Attention-Driven Terminology Probability Estimation for Robust Speech-to-text System

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5189 字 阅读 →
论文解读

Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding

语音编码器 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4347 字 阅读 →
论文解读

Audio-Conditioned Diffusion LLMs for ASR and Deliberation Processing

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4086 字 阅读 →
论文解读

Behind the Scenes: Mechanistic Interpretability of Lora-Adapted Whisper for Speech Emotion Recognition

语音情感识别 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3951 字 阅读 →
论文解读

Benchmarking Humans And Machines On Complex Multilingual Speech Understanding Tasks

音频问答 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3566 字 阅读 →
论文解读

CCST: Cross-Modal and Consistency-Aware Self-Training for Source-Free Unsupervised Domain Adaptation in Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6342 字 阅读 →
论文解读

Contextual Biasing for ASR in Speech LLM with Common Word Cues and Bias Word Position Prediction

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5305 字 阅读 →
论文解读

Cross-Lingual Alzheimer’s Disease Detection with Multimodal LLMs via Speech Cue-Augmented Prompting and Instruction Tuning

语音生物标志物 | 6.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4634 字 阅读 →
论文解读

Cross-Lingual Interleaving for Speech Language Models

语音大模型 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5122 字 阅读 →
论文解读

Cross-Modal Knowledge Distillation for Speech Large Language Models

语音大模型 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3931 字 阅读 →
论文解读

Direct Simultaneous Translation Activation for Large Audio-Language Models

语音翻译 | 6.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4476 字 阅读 →
论文解读

Do Bias Benchmarks Generalise? Evidence from Voice-Based Evaluation of Gender Bias in Speechllms

模型评估 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4695 字 阅读 →