论文解读

UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice

语音翻译 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5273 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-02

共分析 4 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 15 分钟 · 7026 字 阅读 →
论文解读

AppTek Call-Center Dialogues: A Multi-Accent Long-Form Benchmark for English ASR

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4392 字 阅读 →
论文解读

HATS: An Open data set Integrating Human Perception Applied to the Evaluation of Automatic Speech Recognition Metrics

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3963 字 阅读 →
论文解读

Listening with Time: Precise Temporal Awareness for Long-Form Audio Understanding

音频场景理解 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4835 字 阅读 →
论文解读

LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition

语音识别 | 9.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4128 字 阅读 →
论文解读

Spectrographic Portamento Gradient Analysis: A Quantitative Method for Historical Cello Recordings with Application to Beethoven's Piano and Cello Sonatas, 1930--2012

音乐信息检索 | 7.5/10

 · 更新于 2026-09-25 · 约 7 分钟 · 3274 字 阅读 →
论文解读

A Toolkit for Detecting Spurious Correlations in Speech Datasets

模型评估 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4619 字 阅读 →
论文解读

EmoTransCap: Dataset and Pipeline for Emotion Transition-Aware Speech Captioning in Discourses

语音情感识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5250 字 阅读 →
论文解读

StarDrinks: An English and Korean Test Set for SLU Evaluation in a Drink Ordering Scenario

数据集 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3921 字 阅读 →
论文解读

3D Mesh Grid Room Impulse Responses Measured with A Linear Microphone Array And Suppression of Frame Reflections

空间音频 | 8.3/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3526 字 阅读 →
论文解读

A Dataset of Robot-Patient and Doctor-Patient Medical Dialogues for Spoken Language Processing Tasks

语音对话系统 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3628 字 阅读 →
论文解读

A New Method and Dataset for Classroom Teaching Stage Segmentation

课堂阶段分割 | 6.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3742 字 阅读 →
论文解读

A Study of Data Selection Strategies for Pre-Training Self-Supervised Speech Models

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4254 字 阅读 →
论文解读

ACAVCaps: Enabling Large-Scale Training for Fine-Grained and Diverse Audio Understanding

音频分类 | 8.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3768 字 阅读 →
论文解读

AI-Generated Music Detection in Broadcast Monitoring

音频深度伪造检测 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4218 字 阅读 →
论文解读

AISHELL6-Whisper: A Chinese Mandarin Audio-Visual Whisper Speech Dataset with Speech Recognition Baselines

语音识别 | 8.3/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4735 字 阅读 →
论文解读

Aligning Language Models for Lyric-to-Melody Generation with Rule-Based Musical Constraints

音乐生成 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4356 字 阅读 →
论文解读

AMBISONIC-DML: A Benchmark Dataset for Dynamic Higher-Order Ambisonics Music with Motion-Aligned Stems

数据集 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4335 字 阅读 →
论文解读

AnimalCLAP: Taxonomy-Aware Language-Audio Pretraining for Species Recognition and Trait Inference

音频分类 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4960 字 阅读 →