论文解读

Assessing the Impact of Speaker Identity in Speech Spoofing Detection

音频深度伪造检测 | 8.0/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4764 字 阅读 →
论文解读

Assessing The Perceptual Impact of Low-Altitude Aircraft Noise in Cities: An Auralization Framework Using Gaussian Beam Tracing

音频生成 | 8.0/10

 · 更新于 2026-09-11 · 约 8 分钟 · 3778 字 阅读 →
论文解读

Asynchrony-Aware Decoupled Multimodal Control for Cued Speech Video Generation

语音合成 | 7.5/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4846 字 阅读 →
论文解读

ATOM: Adaptive Token-Level Optimal Transport Mixup for Speech Translation

语音翻译 | 8.0/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4810 字 阅读 →
论文解读

Atomic Norm Minimization Revisited: Progressive Atom Identification And Refinement

声源定位 | 7.5/10

 · 更新于 2026-09-11 · 约 8 分钟 · 3620 字 阅读 →
论文解读

Attention-Based Encoder-Decoder Target-Speaker Voice Activity Detection for Robust Speaker Diarization

说话人分离 | 8.0/10

 · 更新于 2026-09-11 · 约 11 分钟 · 5395 字 阅读 →
论文解读

Attention-Weighted Centered Kernel Alignment for Knowledge Distillation in Large Audio-Language Models Applied To Speech Emotion Recognition

语音情感识别 | 8.0/10

 · 更新于 2026-09-11 · 约 12 分钟 · 5893 字 阅读 →
论文解读

Attention2Probability: Attention-Driven Terminology Probability Estimation for Robust Speech-to-text System

语音识别 | 7.0/10

 · 更新于 2026-09-11 · 约 11 分钟 · 5189 字 阅读 →
论文解读

Attentive AV-Fusionnet: Audio-Visual Quality Prediction with Hybrid Attention

音视频 | 7.0/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4767 字 阅读 →
论文解读

Attentive Masked Self-Distillation for Respiratory Sound Classification

音频分类 | 7.5/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4934 字 阅读 →
论文解读

Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding

语音编码器 | 7.5/10

 · 更新于 2026-09-11 · 约 9 分钟 · 4347 字 阅读 →
论文解读

Audience-Aware Co-speech Gesture Generation in Public Speaking via Anticipation Tokens

音频生成 | 8.0/10

 · 更新于 2026-09-11 · 约 11 分钟 · 5224 字 阅读 →
论文解读

Audio Classification Models are Vulnerable to Filter Perturbations

音频分类 | 7.5/10

 · 更新于 2026-09-11 · 约 7 分钟 · 3318 字 阅读 →
论文解读

Audio Deepfake Detection at the First Greeting: "Hi!"

音频深度伪造检测 | 7.5/10

 · 更新于 2026-09-11 · 约 9 分钟 · 4507 字 阅读 →
论文解读

Audio Effect Estimation with DNN-Based Prediction and Search Algorithm

音频效果估计 | 7.0/10

 · 更新于 2026-09-11 · 约 9 分钟 · 4017 字 阅读 →
论文解读

Audio-Conditioned Diffusion LLMs for ASR and Deliberation Processing

语音识别 | 7.0/10

 · 更新于 2026-09-11 · 约 9 分钟 · 4086 字 阅读 →
论文解读

Audio-Guided Multimodal Approach for Fine-Grained Alignment and Boundary Modeling in Active Speaker Detection

说话人检测 | 7.5/10

 · 更新于 2026-09-11 · 约 9 分钟 · 4161 字 阅读 →
论文解读

Audio-Text Jailbreak Attack on Large Audio-Language Models: Towards Generality and Stealthiness

音频安全 | 7.0/10

 · 更新于 2026-09-11 · 约 8 分钟 · 3636 字 阅读 →
论文解读

Audio-to-Score Jazz Solo Transcription with the Rhythm Perceiver

音乐信息检索 | 7.5/10

 · 更新于 2026-09-11 · 约 9 分钟 · 4493 字 阅读 →
论文解读

Audio-Visual Deepfake Generation and Detection: An Exploratory Survey

音频深度伪造检测 | 6.5/10

 · 更新于 2026-09-11 · 约 7 分钟 · 3233 字 阅读 →