论文解读

APKD: Aligned And Paced Knowledge Distillation Towards Lightweight Heterogeneous Multimodal Emotion Recognition

情感识别 | 7.5/10

 · 更新于 2026-09-24 · 约 8 分钟 · 3585 字 阅读 →
论文解读

AQUA-Bench: Beyond finding answers to knowing when there are None in Audio Question Answering

音频问答 | 7.0/10

 · 更新于 2026-09-24 · 约 8 分钟 · 3971 字 阅读 →
论文解读

AR-BSNet: Towards Ultra-Low Complexity Autoregressive Target Speaker Extraction With Band-Split Modeling

语音分离 | 7.0/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4535 字 阅读 →
论文解读

AR&D: A Framework for Retrieving and Describing Concepts for Interpreting AudioLLMs

音频大模型 | 6.5/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4439 字 阅读 →
论文解读

Ara-BEST-RQ: Multi Dialectal Arabic SSL

语音识别 | 6.5/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4587 字 阅读 →
论文解读

Arbitrarily Settable Frame Rate Neural Speech Codec with Content Adaptive Variable Length Segmentation

音频生成 | 7.0/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5159 字 阅读 →
论文解读

ARCHI-TTS: A Flow-Matching-Based Text-to-Speech Model with Self-Supervised Semantic Aligner and Accelerated Inference

语音合成 | 8.0/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6329 字 阅读 →
论文解读

Are Modern Speech Enhancement Systems Vulnerable to Adversarial Attacks?

语音增强 | 7.5/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4629 字 阅读 →
论文解读

Assessing Identity Leakage in Talking Face Generation: Metrics and Evaluation Framework

说话人脸生成 | 7.5/10

 · 更新于 2026-09-24 · 约 8 分钟 · 3595 字 阅读 →
论文解读

Assessing the Impact of Speaker Identity in Speech Spoofing Detection

音频深度伪造检测 | 8.0/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4764 字 阅读 →
论文解读

Assessing The Perceptual Impact of Low-Altitude Aircraft Noise in Cities: An Auralization Framework Using Gaussian Beam Tracing

音频生成 | 8.0/10

 · 更新于 2026-09-24 · 约 8 分钟 · 3778 字 阅读 →
论文解读

Asynchrony-Aware Decoupled Multimodal Control for Cued Speech Video Generation

语音合成 | 7.5/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4846 字 阅读 →
论文解读

ATOM: Adaptive Token-Level Optimal Transport Mixup for Speech Translation

语音翻译 | 8.0/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4810 字 阅读 →
论文解读

Atomic Norm Minimization Revisited: Progressive Atom Identification And Refinement

声源定位 | 7.5/10

 · 更新于 2026-09-24 · 约 8 分钟 · 3620 字 阅读 →
论文解读

Attention-Based Encoder-Decoder Target-Speaker Voice Activity Detection for Robust Speaker Diarization

说话人分离 | 8.0/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5395 字 阅读 →
论文解读

Attention-Weighted Centered Kernel Alignment for Knowledge Distillation in Large Audio-Language Models Applied To Speech Emotion Recognition

语音情感识别 | 8.0/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5893 字 阅读 →
论文解读

Attention2Probability: Attention-Driven Terminology Probability Estimation for Robust Speech-to-text System

语音识别 | 7.0/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5189 字 阅读 →
论文解读

Attentive AV-Fusionnet: Audio-Visual Quality Prediction with Hybrid Attention

音视频 | 7.0/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4767 字 阅读 →
论文解读

Attentive Masked Self-Distillation for Respiratory Sound Classification

音频分类 | 7.5/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4934 字 阅读 →
论文解读

Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding

语音编码器 | 7.5/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4347 字 阅读 →