论文解读

Improving Anomalous Sound Detection with Attribute-Aware Representation from Domain-Adaptive Pre-Training

音频事件检测 | 8.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5175 字 阅读 →
论文解读

Improving Audio Event Recognition with Consistency Regularization

音频事件检测 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4143 字 阅读 →
论文解读

Input-Adaptive Differentiable Filterbanks via Hypernetworks for Robust Speech Processing

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4225 字 阅读 →
论文解读

Is Phase Really Needed for Weakly-Supervised Dereverberation?

语音增强 | 6.0/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3528 字 阅读 →
论文解读

KAN We Make Models Simpler for Audio Deepfake Detection with Kolmogorov–Arnold Networks?

音频深度伪造检测 | 7.5/10

 · 更新于 2026-09-06 · 约 7 分钟 · 3362 字 阅读 →
论文解读

Leveraging Audio-Visual Data to Reduce the Multilingual Gap in Self-Supervised Speech Models

语音识别 | 6.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4691 字 阅读 →
论文解读

Leveraging Segment-Level Speech Representations for LLM-Based Speech Recognition

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5225 字 阅读 →
论文解读

Lightweight and Perceptually-Guided Voice Conversion for Electro-Laryngeal Speech

语音转换 | 7.5/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6301 字 阅读 →
论文解读

Localizing Speech Deepfakes Beyond Transitions via Segment-Aware Learning

音频深度伪造检测 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4261 字 阅读 →
论文解读

Matrix-Structured Hierarchical Convolutional Modeling for Pronunciation Assessment and Mispronunciation Detection

语音评估 | 8.0/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6402 字 阅读 →
论文解读

Measuring Prosody Diversity in Zero-Shot TTS: A New Metric, Benchmark, and Exploration

语音合成 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4416 字 阅读 →
论文解读

Mind the Shift: Using Delta SSL Embeddings to Enhance Child ASR

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3909 字 阅读 →
论文解读

MR-FlowDPO: Multi-Reward Direct Preference Optimization for Flow-Matching Text-to-Music Generation

音乐生成 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5357 字 阅读 →
论文解读

MT-HuBERT: Self-Supervised Mix-Training for Few-Shot Keyword Spotting in Mixed Speech

关键词检测 | 7.0/10

 · 更新于 2026-09-06 · 约 23 分钟 · 11472 字 阅读 →
论文解读

Multi-Layer Attentive Probing Improves Transfer of Audio Representations for Bioacoustics

生物声学 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4488 字 阅读 →
论文解读

Multi-Scale Physiologically-Motivated Alignment for Auditory Attention Decoding

听觉注意力解码 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4347 字 阅读 →
论文解读

Multi-View Hierarchical Hypergraph Neural Network for Automatic Stuttering Detection

语音生物标志物 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5262 字 阅读 →
论文解读

On deepfake voice detection - It’s all in the presentation

音频深度伪造检测 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4174 字 阅读 →
论文解读

Online Register For Dual-Mode Self-Supervised Speech Models: Mitigating the Lack of Future Context

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5121 字 阅读 →
论文解读

Optimizing Domain-Adaptive Self-Supervised Learning for Clinical Voice-Based Disease Classification

语音生物标志物 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4984 字 阅读 →