论文解读

RASD-SR: A Robust Anomalous Sound Detection Framework with Score Recalibration

异常声音检测 | 8.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5155 字 阅读 →
论文解读

Reading Between the Waves: Robust Topic Segmentation Using Inter-Sentence Audio Features

音频分类 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5113 字 阅读 →
论文解读

Reasoning Driven Captions to Assist Noise Robust Speech Emotion Recognition

语音情感识别 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4690 字 阅读 →
论文解读

Reducing Prompt Sensitivity in LLM-Based Speech Recognition Through Learnable Projection

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4818 字 阅读 →
论文解读

Regularized Inverse Filter Design for Rigid Spherical Microphone Array Processing: Laplace- And Time-Domain Representations

空间音频 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4068 字 阅读 →
论文解读

RMODGDF: A Robust STFT-Derived Feature for Musical Instrument Recognition

音乐信息检索 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4290 字 阅读 →
论文解读

Robust and Lightweight F0 Estimation Through Mid-Level Fusion of DSP-Informed Features

基频估计 | 8.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5195 字 阅读 →
论文解读

Robust Deepfake Audio Detection via Multi-Level Intermediate Feature Fusion

音频深度伪造检测 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3781 字 阅读 →
论文解读

RoCo: Robust Code for Fast and Effective Proactive Defense against Voice Cloning Attack

音频安全 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5166 字 阅读 →
论文解读

RRPO: Robust Reward Policy Optimization for LLM-Based Emotional TTS

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4067 字 阅读 →
论文解读

Sampling-Rate-Agnostic Speech Super-Resolution Based on Gaussian Process Dynamical Systems with Deep Kernel Learning

语音增强 | 6.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4782 字 阅读 →
论文解读

Snore Sound Classification Based on Physiological Features and Adaptive Loss Function

音频分类 | 6.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5116 字 阅读 →
论文解读

Spectral or Spatial? Leveraging Both for Speaker Extraction in Challenging Data Conditions

语音分离 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5058 字 阅读 →
论文解读

Spectrogram Event Based Feature Representation for Generalizable Automatic Music Transcription

音乐信息检索 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 5000 字 阅读 →
论文解读

Spiking Attention Network: A Hybrid Neuromorphic Approach to Underwater Acoustic Localization and Zero-Shot Adaptation

声源定位 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3981 字 阅读 →
论文解读

Staged Diffusion with Hybrid Mixture-of-Experts (MOE) for Multimodal Sentiment Analysis

语音情感识别 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4137 字 阅读 →
论文解读

StreamMark: A Deep Learning-Based Semi-Fragile Audio Watermarking for Proactive Deepfake Detection

音频深度伪造检测 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4723 字 阅读 →
论文解读

SURE: Synergistic Uncertainty-Aware Reasoning for Multimodal Emotion Recognition in Conversations

语音情感识别 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4997 字 阅读 →
论文解读

Target-Speaker LLM-ASR with Speaker-Aware Speech Encoder

语音识别 | 8.8/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4444 字 阅读 →
论文解读

Toward Robust And Efficient Beat Tracking Via Beat-Aware Attention

音乐理解 | 8.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4385 字 阅读 →