论文解读

Speech Entrainment in Multi-Party Conversations with a Digital Agent

语音交互 | 5.3/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5599 字 阅读 →
论文解读

Towards High-Level Semantic Intelligence

音视频理解 | 7.0/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6680 字 阅读 →
论文解读

An Evaluation Framework for Structured Audio Captions Validated by Controlled Perturbations

数据集 | 5.3/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7447 字 阅读 →
论文解读

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion

语音转换 | 7.9/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5983 字 阅读 →
论文解读

Validating the Single Item Kawaii Measure

音频理解 | 6.4/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8270 字 阅读 →
论文解读

Summary of DCASE 2026 Task 5: Audio-Dependent Question Answering

音频理解 | 5.4/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7727 字 阅读 →
论文解读

FillGauss: Fine-Grained Filling-Aware Impact Sound Generation for 3D Gaussian Splatting

音频生成 | 6.2/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6266 字 阅读 →
论文解读

RealDESED: A Real-World Domestic Sound Event Detection Benchmark

音频事件检测 | 7.9/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7315 字 阅读 →
论文解读

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos

音视频理解 | 6.0/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7850 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-07-20

共分析 15 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 52 分钟 · 25690 字 阅读 →
论文解读

InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring

多模态模型 | 7.3/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5751 字 阅读 →
论文解读

MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation

音视频生成 | 6.3/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6689 字 阅读 →
论文解读

SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings

语音分离 | 8.1/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8674 字 阅读 →
论文解读

From Continuous Deployment to Queryable Dataset: Terabyte-Scale AIS-Aligned Passive Acoustic Labelling

音频理解 | 6.1/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6337 字 阅读 →
论文解读

Self-supervised Speech Comparison for L2 Phone, Rhythm, and Intonation Scoring

音频理解 | 7.7/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6639 字 阅读 →
论文解读

VIP-MINGLE: A Corpus for Videoconference and In-Person Multimodal Interaction in Group Language Engagement

音频理解 | 6.5/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7416 字 阅读 →
论文解读

An Objective Intelligibility Metric Evaluation on Spanish Speech

语音质量评估 | 6.2/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6711 字 阅读 →
论文解读

VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion

语音伪造检测 | 8.1/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6950 字 阅读 →
论文解读

Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks

多模态模型 | 6.9/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6518 字 阅读 →
论文解读

Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks

多模态模型 | 6.3/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5232 字 阅读 →