论文解读

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling

语音识别 | 8.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5701 字 阅读 →
论文解读

ParaBridge: Bridging Paralinguistic Perception and Dialogue Behavior in Speech Language Models

参数高效微调 | 8.6/10

 · 更新于 2026-09-25 · 约 6 分钟 · 2708 字 阅读 →
论文解读

Phoneme-First Prediction for LLM-Based Speech Recognition

语音识别 | 6.9/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6037 字 阅读 →
论文解读

RespiraMFM: A Multimodal Foundation Model with Contrastive Audio-Language Alignment for Respiratory Disease Identification

对比学习 | 6.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6068 字 阅读 →
论文解读

Speech Encoder Fusion for LLM-based Automatic Speech Recognition

语音识别 | 7.2/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5281 字 阅读 →
论文解读

Towards Deep Contextual Reasoning from Broad Descriptions for ASR with Speech-LLM via Metadata-Driven Reasoning Chains

语音识别 | 6.2/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4631 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-10

共分析 45 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 126 分钟 · 62690 字 阅读 →
论文解读

Rethinking Depth: A study of the Recursive-Transformer for Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5625 字 阅读 →
论文解读

Subtitle-Aligned Fine-Tuning of Whisper for Swiss German ASR: Benchmark Contamination, Convention Mismatch, and an Honest Baseline at 25.6% WER (13.8% cWER)

语音识别 | 8.2/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6640 字 阅读 →
论文解读

TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech

语音合成 | 8.2/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6310 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-09

共分析 48 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 141 分钟 · 70154 字 阅读 →
论文解读

Acoustic Cue Alignment in Audio Language Models for Speech Emotion Recognition

语音情感识别 | 7.8/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5285 字 阅读 →
论文解读

Beyond Semantic Dominance: Cognitive Affective Reasoning and Empathetic Response Alignment in Audio Language Models

语音合成 | 9.2/10

 · 更新于 2026-09-25 · 约 26 分钟 · 12718 字 阅读 →
论文解读

How Far Can Chord-Symbol Time-Series Adaptation Carry Genre Identity? Capabilities and Boundaries in Multi-Genre Chord-Symbol Modeling

音乐信息检索 | 8.8/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5729 字 阅读 →
论文解读

KIT's Submission to Cross-Lingual Voice Cloning in IWSLT 2026

语音合成 | 7.2/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5494 字 阅读 →
论文解读

SpectCount: Spectrotemporal Counting via Synthetic Signals Improves Large Audio Language Models

数据增强 | 5.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4925 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-08

共分析 38 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 108 分钟 · 53770 字 阅读 →
论文解读

Age-Aware Adapter Tuning for Children's Speech Recognition

语音识别 | 8.4/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5531 字 阅读 →
论文解读

Beyond Generative Decoding: Discriminative Hidden-State Readout from a Native Omni-Modal LLM for Multimodal Sentiment Analysis

多模态模型 | 5.3/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6789 字 阅读 →
论文解读

FiLM-Based Speaker Conditioning of a SpeechLLM for Pathological Speech Recognition

语音识别 | 8.0/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7310 字 阅读 →