论文解读

Co-policy: Responsive Human-Robot Co-Creation for Musical Performances

音乐生成 | 8.5/10

 · 更新于 2026-09-07 · 约 14 分钟 · 6648 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-22

共分析 1 篇语音/AI 论文

 · 更新于 2026-09-07 · 约 4 分钟 · 1684 字 阅读 →
论文解读

A Comparative Study of Pretrained Transformer Models for Quranic ASR: Speech Representations, Label Formats, and Dataset Composition

语音识别 | 7.2/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5716 字 阅读 →
论文解读

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine

语音合成 | 7.4/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5736 字 阅读 →
论文解读

Analyzing Language and Geographical Variation in Speech Representations Across 60 Indic Languages

语音识别 | 6.5/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5299 字 阅读 →
论文解读

Beyond Speaker Independence: Evaluating Cross-Lingual Acoustic-to-Articulatory Inversion Across Finnish and Russian

自监督学习 | 4.9/10

 · 更新于 2026-09-07 · 约 14 分钟 · 6597 字 阅读 →
论文解读

Cross-Dataset, Age, and Gender Generalization: A Comprehensive Analysis of Fine-Tuning Strategies for Low-Resource Children's ASR

语音识别 | 6.7/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5701 字 阅读 →
论文解读

Exploring Feature Extraction Technique Parameters for Acoustic Gunshot Classification

音频事件检测 | 7.9/10

 · 更新于 2026-09-07 · 约 10 分钟 · 4882 字 阅读 →
论文解读

Exploring Pre-training Benefits on Phoneme Addition through Fine-tuning in Speech Synthesis

Exploring Pre-training Benefits on Phoneme Addition through Fine-tuning in Speech Synthesis

 · 更新于 2026-09-07 · 约 21 分钟 · 10509 字 阅读 →
论文解读

FlowEdit: Associative Memory for Lifelong Pronunciation Adaptation in Flow-Matching TTS

语音合成 | 10/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5749 字 阅读 →
论文解读

FlowFake: Liquid Networks for Audio Deepfake Detection

模型压缩 | 8.5/10

 · 更新于 2026-09-07 · 约 14 分钟 · 6691 字 阅读 →
论文解读

How Do Instructions Shape Speech? Cross-Attention Attribution for Style-Captioned Text-to-Speech

语音合成 | 7.7/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5246 字 阅读 →
论文解读

Hybrid Diffusion Transformer for Instruction-Guided Audio Editing via Rectified Flow

Transformer | 7.6/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5452 字 阅读 →
论文解读

IHBench: Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows

语音对话系统 | 7.5/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5988 字 阅读 →
论文解读

Improving Code-Switching ASR with Code-Mixing Guided Synthetic Speech

语音识别 | 7.6/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5725 字 阅读 →
论文解读

Improving End-to-End Speech Recognition for Dysarthric Speech through In-Domain Data Augmentation

语音识别 | 6.5/10

 · 更新于 2026-09-07 · 约 23 分钟 · 11077 字 阅读 →
论文解读

Interpreting Content and Speaker Characteristics in Factorised Self-Supervised Subspaces

语音合成 | 5/10

 · 更新于 2026-09-07 · 约 10 分钟 · 4928 字 阅读 →
论文解读

Investigating Human-Model Discrepancies in Speech Quality Assessment via Acoustic and Prosodic Perturbations

语音合成 | 6.9/10

 · 更新于 2026-09-07 · 约 16 分钟 · 7804 字 阅读 →
论文解读

Latency-Configurable Streaming Speech Enhancement via Asymmetric Temporal Padding

语音增强 | 7.2/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5381 字 阅读 →
论文解读

Leveraging systems' non-linearity to tackle the scarcity of data in the design of Intelligent Fault Diagnosis Systems

数据增强 | 5.5/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5424 字 阅读 →