论文解读

Mitigating Over-Suppression in Speech Enhancement via Inference-Time Rethink-and-Refine Correction Module

语音增强 | 6.2/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5718 字 阅读 →
论文解读

Physics-Informed Learning for Robust Acoustic Localization with Calibrated Uncertainty

声源定位 | 6.2/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6273 字 阅读 →
论文解读

Assessing AI-generated music detection in real-world broadcast monitoring

音频伪造检测 | 6.8/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6180 字 阅读 →
论文解读

FormBharo: Designing and Evaluating a Voice Agent for Conversational Form Filling in Rural India

语音交互 | 7.0/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6937 字 阅读 →
论文解读

How to Recognize New Words: A Comparison Between Context Biasing Methods and Speech LLMs

语音识别 | 6.4/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7408 字 阅读 →
论文解读

Deep Learning for Real-Time Sound Order Recognition in Human-Robot Interaction

音频事件检测 | 5.2/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5506 字 阅读 →
论文解读

Smartphone Audio Based Distress Detection

音频事件检测 | 6.3/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7186 字 阅读 →
论文解读

StuPASE: Towards Low-Hallucination Studio-Quality Generative Speech Enhancement

语音增强 | 7.0/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7515 字 阅读 →
论文解读

Rethinking Modality Reliability in Multimodal Sentiment Analysis with Incomplete Observations

音视频理解 | 6.2/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7260 字 阅读 →
论文解读

Leveraging Beam Search Information for Confidence Estimation in E2E ASR

语音识别 | 6.3/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7632 字 阅读 →
论文解读

Improved Robustness in AI-Generated Music Detection

音频伪造检测 | 8.0/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7579 字 阅读 →
论文解读

Teffic-Audio: Tell Fact from Fiction

语音伪造检测 | 6.8/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8390 字 阅读 →
论文解读

Voice Memory for Agentic Speech Recognition

语音识别 | 8.2/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8134 字 阅读 →
论文解读

Automatic Audio Equalization with Semantic Embeddings

语音增强 | 5.8/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7531 字 阅读 →
论文解读

Investigating Codec-Internal Latent Audio Watermarking for Neural Codec Robustness

音频水印 | 6.4/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7234 字 阅读 →
论文解读

Towards Array-Invariant Speech Enhancement via Geometry-Aware Dynamic Convolution

语音增强 | 6.3/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6187 字 阅读 →
论文解读

How Reliable Are Multimodal Signals of Conversational State? Evidence from Remote Dyadic Collaborative Tasks

鲁棒性 | 6.6/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7971 字 阅读 →
论文解读

Time-Frequency Consistency Learning for Robust Speech Deepfake Detection

语音伪造检测 | 7.9/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6790 字 阅读 →
论文解读

Controlling Implicit Shortcut Reliance in L2 Spoken English Auto-markers

语音质量评估 | 7.5/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8422 字 阅读 →
论文解读

Natural Backdoor Attacks on Speech Recognition Models

语音识别 | 3.5/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7287 字 阅读 →