论文解读

LuxEmo: Expressive Text-to-Speech Corpus for Luxembourgish

语音合成 | 7.5/10

 · 更新于 2026-09-07 · 约 13 分钟 · 6067 字 阅读 →
论文解读

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs

语音合成 | 7.9/10

 · 更新于 2026-09-07 · 约 13 分钟 · 6024 字 阅读 →
论文解读

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation

语音识别 | 7/10

 · 更新于 2026-09-07 · 约 6 分钟 · 2889 字 阅读 →
论文解读

Probing-Guided Layer Selection from Self-Supervised Speech Models for Generalizable Audio Deepfake Detection

集成学习 | 7.5/10

 · 更新于 2026-09-07 · 约 13 分钟 · 6363 字 阅读 →
论文解读

Reference-Based Prosody and Rhythm Evaluation for Spoken Dialogue Systems

语音对话系统 | 4.7/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5571 字 阅读 →
论文解读

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation

知识蒸馏 | 10/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5503 字 阅读 →
论文解读

SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation

语音合成 | 7.5/10

 · 更新于 2026-09-07 · 约 14 分钟 · 6593 字 阅读 →
论文解读

Tone-Conditioned Curriculum Learning for Low-Resource Bantu Speech Recognition

语音识别 | 7.3/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5903 字 阅读 →
论文解读

UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling

语音合成 | 7.3/10

 · 更新于 2026-09-07 · 约 5 分钟 · 2255 字 阅读 →
论文解读

What Counts as an Error? Dual-Reference Benchmarking for Atypical ASR

语音识别 | 7.3/10

 · 更新于 2026-09-07 · 约 25 分钟 · 12422 字 阅读 →
论文解读

ZEBRA: Zero-Shot Entropy-Regularized Prompt Learning for Base-to-Novel Generalization in Audio-Language Models

音频分类 | 7.1/10

 · 更新于 2026-09-07 · 约 13 分钟 · 6499 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-07-01

共分析 35 篇语音/AI 论文

 · 更新于 2026-09-07 · 约 102 分钟 · 50963 字 阅读 →
论文解读

Agent-Computer Observation Interfaces Enable Dynamic Computer Use

语音识别 | 8.4/10

 · 更新于 2026-09-07 · 约 14 分钟 · 6790 字 阅读 →
论文解读

AMR: Adaptive Modality Routing for Multimodal Polyglot Speaker Identification

说话人识别 | 7.8/10

 · 更新于 2026-09-07 · 约 13 分钟 · 6161 字 阅读 →
论文解读

Child-Centric Voice Anonymization in Single and Multi-Speaker Speech via Domain-Adapted SSL Models

语音匿名化 | 7.2/10

 · 更新于 2026-09-07 · 约 14 分钟 · 6716 字 阅读 →
论文解读

Clustering Unsupervised Representations as Defense against Poisoning Attacks on Speech Commands Classification System

自监督学习 | 6.5/10

 · 更新于 2026-09-07 · 约 16 分钟 · 7576 字 阅读 →
论文解读

Comparing Human and Automatic Recognition of Dutch Dysarthric Continuous Speech: A Case Study

语音识别 | 6.5/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5900 字 阅读 →
论文解读

CTC-Seeded Token Edit Refinement for Non-Autoregressive Speech Recognition

语音识别 | 7.7/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5815 字 阅读 →
论文解读

DialogPII: A multilingual dataset of synthetic dialog transcripts to detect personal information

语音识别 | 8.9/10

 · 更新于 2026-09-07 · 约 35 分钟 · 17529 字 阅读 →
论文解读

DTM-Codec: Dynamic Token Masking for VFR Speech Coding with Efficient Boundary Selection

语音编码 | 8.1/10

 · 更新于 2026-09-07 · 约 18 分钟 · 8901 字 阅读 →