论文解读

An Acoustic Landmark Database of the English Lexicon via Articulatory Synthesis

语音合成 | 6.9/10

 · 更新于 2026-09-07 · 约 16 分钟 · 7940 字 阅读 →
论文解读

An Analysis of Untrained Deep Reservoir Networks for Audio Surveillance

音频事件检测 | 8.8/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5526 字 阅读 →
论文解读

An Evaluation Framework for Text-to-Speech Voice Reconstruction

语音合成 | 8.8/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5642 字 阅读 →
论文解读

An implicitization-based solution to the minimal 4s/6r ToA problem using Cayley--Menger determinants

An implicitization-based solution to the minimal 4s/6r ToA problem using Cayley--Menger determinants

 · 更新于 2026-09-07 · 约 11 分钟 · 5023 字 阅读 →
论文解读

AOR-Bench: Do Large Audio Language Models Over-Refuse Pseudo-Harmful Queries?

音频问答 | 7.9/10

 · 更新于 2026-09-07 · 约 10 分钟 · 4940 字 阅读 →
论文解读

ATCCaps: A Call-Sign-Aware Speech Dataset for Air Traffic Control Recognition

语音识别 | 8.6/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5325 字 阅读 →
论文解读

Audio Editing in the Era of Foundation Models: A Survey

Audio Editing in the Era of Foundation Models: A Survey

 · 更新于 2026-09-07 · 约 12 分钟 · 5598 字 阅读 →
论文解读

AudioCALM: Continuous Autoregressive Language Modeling for Universal Audio Generation

语音合成 | 7.9/10

 · 更新于 2026-09-07 · 约 15 分钟 · 7085 字 阅读 →
论文解读

AugCodec: A Low-Bitrate Disentangled Neural Speech Codec via Data Augmentation

数据增强 | 6.7/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5361 字 阅读 →
论文解读

Backdoor Attacks on Speech Emotion Recognition via TTS-Generated Poisoning

语音情感识别 | 7/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5489 字 阅读 →
论文解读

Bagpiper-Edit: Zero-Shot Open-Ended Audio Editing via Rich-Caption

语音合成 | 7.6/10

 · 更新于 2026-09-07 · 约 14 分钟 · 6835 字 阅读 →
论文解读

Bagpiper-TTS: Natural Language Guided Universal Speech Synthesis

语音合成 | 8.4/10

 · 更新于 2026-09-07 · 约 22 分钟 · 10533 字 阅读 →
论文解读

Benchmarking Large Language Models for Grapheme-to-Phoneme Conversion: A Japanese Case Study

语音合成 | 8.4/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5133 字 阅读 →
论文解读

Beyond ROC-AUC: Operating-Point Performance Reporting for Biometric Verification

Beyond ROC-AUC: Operating-Point Performance Reporting for Biometric Verification

 · 更新于 2026-09-07 · 约 11 分钟 · 5277 字 阅读 →
论文解读

Bridging Self-Supervised Learning and Speech Enhancement: A Wav2Vec2-Conditioned Framework

语音增强 | 7.5/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5269 字 阅读 →
论文解读

Bridging the Age Gap: Towards Detecting Neural Audio Codec Synthesized Elderly Speech Deepfake

语音伪造检测 | 8.6/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5580 字 阅读 →
论文解读

CAAD: Contrastive Audio-Aware Distillation for Efficient Speech Language Models

语音识别 | 8.9/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5368 字 阅读 →
论文解读

CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales

语音识别 | 8.6/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5766 字 阅读 →
论文解读

Catching Lies Without Sending the Video: Privacy-Preserving Multimodal Deception Detection

多模态模型 | 6.2/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5674 字 阅读 →
论文解读

Compiling Differentiable Audio Graphs to Real-Time DSP

Compiling Differentiable Audio Graphs to Real-Time DSP

 · 更新于 2026-09-07 · 约 12 分钟 · 5728 字 阅读 →