论文解读

Modeling Stylistic Co-evolution in Symbolic Music Heritage Collections

音乐理解 | 4.9/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8057 字 阅读 →
论文解读

Multimodal Domain Generalization for Depression Detection: An Attention-Based BiLSTM Network with Domain-Adversarial Training

音频分类 | 6.4/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6423 字 阅读 →
论文解读

PathRIR: Physics-Guided Acoustic Path Selection and Late-Tail Compensation for Fast Room Impulse Response Simulation

空间音频 | 7.2/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6462 字 阅读 →
论文解读

Resource-Aware Topology Management for ISAC-Enabled TDOA Localization in IoUT Networks

声源定位 | 4.1/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5538 字 阅读 →
论文解读

Revisiting Vocos: That Phasiness Business in Time-Frequency Neural Vocoding

语音合成 | 7.1/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6739 字 阅读 →
论文解读

Speech Entrainment in Multi-Party Conversations with a Digital Agent

语音交互 | 5.3/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5599 字 阅读 →
论文解读

Speech Signals Complement LLMs for Predicting Interpersonal Attraction in Speed Dating

语音交互 | 6.9/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7083 字 阅读 →
论文解读

How Meta-Learning Shapes LoRA Adapter Geometry in Speech Deepfake Detection

语音伪造检测 | 5.2/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7098 字 阅读 →
论文解读

MemNMF: Memory-Augmented NMF on LPC Spectra for Anomalous Sound Detection

音频事件检测 | 5.3/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6236 字 阅读 →
论文解读

Music-JEPA: Learning a World Model of Sound from Action

音乐转录 | 6.2/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6393 字 阅读 →
论文解读

Phylogenetic signal in marine mammal and bird vocalizations captured by audio foundation models: the limited benefit of domain-specific pretraining

音频理解 | 7.7/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6153 字 阅读 →
论文解读

Probing Speaker Identity Sensitivity in Audio Deepfake Detectors

语音伪造检测 | 6.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6060 字 阅读 →
论文解读

SoundscapeAgent: Agentic Soundscape Construction for Controllable Synthesis and Scalable Audio-Language Supervision

大语言模型 | 8.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5047 字 阅读 →
论文解读

Synthetic Speech, Real Signal: Paralinguistic Preservation and Cross-Lingual Augmentation via Voice Cloning

语音情感识别 | 6.2/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6667 字 阅读 →
论文解读

Transforming Keystroke Noise to Text: Self-Supervised Acoustic Eavesdropping Attacks on Keyboards

语音活动检测 | 6.5/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7374 字 阅读 →
论文解读

From Read Speech to Spoken Digits: A Task-Specific Evaluation of Speech Privacy With Informed Attackers

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9074 字 阅读 →
论文解读

Improving the performance of an ASV system using hybrid speech features

说话人验证 | 5.0/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6562 字 阅读 →
论文解读

Instruct-FD: Can Your Full-Duplex Speech System Follow Turn-Taking Instructions?

语音交互 | 7.2/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6484 字 阅读 →
论文解读

TF-MossFormer: Integrating Convolution Gated Local-Global Attentions for Enhanced Time-Frequency Domain Monaural Speech Separation

语音分离 | 6.3/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5695 字 阅读 →
论文解读

Word meaning co-determines vowel-inherent spectral change. A corpus-based investigation of conversational Mandarin

语音属性识别 | 5.9/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7278 字 阅读 →