论文解读

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation

语音情感识别 模型评估 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5334 字 阅读 →
论文解读

A Bayesian Approach to Singing Skill Evaluation Using Semitone Pitch Histogram and MCMC-Based Generated Quantities

音乐理解 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4110 字 阅读 →
论文解读

A Dataset of Robot-Patient and Doctor-Patient Medical Dialogues for Spoken Language Processing Tasks

语音对话系统 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3628 字 阅读 →
论文解读

A Superb-Style Benchmark of Self-Supervised Speech Models for Audio Deepfake Detection

音频深度伪造检测 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4149 字 阅读 →
论文解读

A Text-To-Text Alignment Algorithm for Better Evaluation of Modern Speech Recognition Systems

模型评估 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3720 字 阅读 →
论文解读

Acoustic Non-Stationarity Objective Assessment with Hard Label Criteria for Supervised Learning Models

音频分类 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3845 字 阅读 →
论文解读

Adaptive Spectral Weighting in Sagittal-Plane Sound Localization: A Reliability-Driven Approach

声源定位 | 6.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4116 字 阅读 →
论文解读

Aligning Generative Speech Enhancement with Perceptual Feedback

语音增强 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5533 字 阅读 →
论文解读

AR&D: A Framework for Retrieving and Describing Concepts for Interpreting AudioLLMs

音频大模型 | 6.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4439 字 阅读 →
论文解读

Assessing Identity Leakage in Talking Face Generation: Metrics and Evaluation Framework

说话人脸生成 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3595 字 阅读 →
论文解读

Attention-Based Encoder-Decoder Target-Speaker Voice Activity Detection for Robust Speaker Diarization

说话人分离 | 8.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5395 字 阅读 →
论文解读

Attentive AV-Fusionnet: Audio-Visual Quality Prediction with Hybrid Attention

音视频 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4767 字 阅读 →
论文解读

Auditory Illusion Benchmark for Large Audio Language Models

模型评估 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3529 字 阅读 →
论文解读

Automatic Estimation of Speaker Diarization Error Rate Based on Features of Audio Quality and Speaker Discriminability

说话人分离 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4339 字 阅读 →
论文解读

AVO-65: A Large-Scale Hierarchical Audio-Visual Object Dataset

音视频 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4186 字 阅读 →
论文解读

Benchmarking Humans And Machines On Complex Multilingual Speech Understanding Tasks

音频问答 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3566 字 阅读 →
论文解读

Benchmarking Music Autotagging with MGPHot Expert Annotations vs. Generic Tag Datasets

音乐信息检索 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4442 字 阅读 →
论文解读

Break-the-Beat! Controllable MIDI-to-Drum audio synthesis

音乐生成 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4419 字 阅读 →
论文解读

BridgeCode: A Dual Speech Representation Paradigm for Autoregressive Zero-Shot Text-to-Speech Synthesis

语音合成 | 8.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5196 字 阅读 →
论文解读

BSMP-SENet:Band-Split Magnitude-Phase Network for Speech Enhancement

语音增强 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4436 字 阅读 →