论文解读

An Acoustic Landmark Database of the English Lexicon via Articulatory Synthesis

语音合成 | 6.9/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7940 字 阅读 →
论文解读

ATCCaps: A Call-Sign-Aware Speech Dataset for Air Traffic Control Recognition

语音识别 | 8.6/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5325 字 阅读 →
论文解读

Bridging the Age Gap: Towards Detecting Neural Audio Codec Synthesized Elderly Speech Deepfake

语音伪造检测 | 8.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5580 字 阅读 →
论文解读

CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales

语音识别 | 8.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5766 字 阅读 →
论文解读

HALAS: A Human-Annotated Dataset of Hallucinations of Modern ASR Systems

语音识别 | 8.4/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5820 字 阅读 →
论文解读

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine

语音合成 | 7.4/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5736 字 阅读 →
论文解读

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models

多模态模型 | 8.1/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5119 字 阅读 →
论文解读

SingFox: A Multi-Lingual Singfake Detection Corpus

语音伪造检测 | 5.4/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6765 字 阅读 →
论文解读

L-Proto: Language-Aware Episodic Prototypical Training for Multilingual Speaker Verification

说话人验证 | 7.1/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5871 字 阅读 →
论文解读

OlfactProfile: Profile-Conditioned Odor Prediction from Audiovisual Content

多模态模型 | 5.6/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5084 字 阅读 →
论文解读

When Multiple Scripts Matter: Evaluating ASR in Clinical Settings

语音识别 | 9.1/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6513 字 阅读 →
论文解读

ROMPAR: Morphological Completion and Demographic Unlearning for Romanian-Accented Speech Recognition

语音识别 | 6.2/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4459 字 阅读 →
论文解读

TuneJury: An Open Metric for Improving Music Generation Preference Alignment

多模态模型 | 9.7/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6483 字 阅读 →
论文解读

HIDVAS: A Hearing Instrument Dataset in Various Acoustical Scenarios for Algorithm Evaluation and Training

语音增强 | 9/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4950 字 阅读 →
论文解读

Spatio-Temporal Audio Language Modeling for Dynamic Sound Sources

音频问答 | 6.9/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5968 字 阅读 →
论文解读

Decoding Insect Song: A Multitask Semisupervised Orthoptera Bioacoustic Classifier

音频分类 | 8.7/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6708 字 阅读 →
论文解读

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents

数据集 | 7/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5813 字 阅读 →
论文解读

I Understand How You Feel: Enhancing Deeper Emotional Support Through Multilingual Emotional Validation in Dialogue System

语音识别 | 5.8/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7227 字 阅读 →
论文解读

Pretrained self-supervised speech models can recognize unseen consonants

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4929 字 阅读 →
论文解读

SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations

语音合成 | 7.9/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4632 字 阅读 →