论文解读

Mind the Microphone Gap: Benchmarking Array Upsampling Strategies for Latent Acoustic Mapping

声源定位 | 6.6/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6635 字 阅读 →
论文解读

Modeling Stylistic Co-evolution in Symbolic Music Heritage Collections

音乐理解 | 4.9/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8057 字 阅读 →
论文解读

MoLGE: Mixture of Language Group Experts for Efficient Scaling of Massively Multilingual Speech Recognition

语音识别 | 6.7/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4906 字 阅读 →
论文解读

Multimodal Domain Generalization for Depression Detection: An Attention-Based BiLSTM Network with Domain-Adversarial Training

音频分类 | 6.4/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6423 字 阅读 →
论文解读

Music-Source-Separation-Training (MSST): A Unified Framework for Training and Evaluating Music Demixing Models

音乐源分离 | 5.7/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7537 字 阅读 →
论文解读

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation

音视频生成 | 7.1/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7731 字 阅读 →
论文解读

PathRIR: Physics-Guided Acoustic Path Selection and Late-Tail Compensation for Fast Room Impulse Response Simulation

空间音频 | 7.2/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6462 字 阅读 →
论文解读

Resource-Aware Topology Management for ISAC-Enabled TDOA Localization in IoUT Networks

声源定位 | 4.1/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5538 字 阅读 →
论文解读

Revisiting Vocos: That Phasiness Business in Time-Frequency Neural Vocoding

语音合成 | 7.1/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6739 字 阅读 →
论文解读

Simple Language Normalization Wins: Cross-Lingual Speaker Verification for the TidyVoice 2026 Challenge

说话人验证 | 6.8/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6547 字 阅读 →
论文解读

Singlish, Can or Not? Fine-Tuning and Evaluating Zero-Shot TTS for Singapore English

语音合成 | 6.3/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7963 字 阅读 →
论文解读

Speech Entrainment in Multi-Party Conversations with a Digital Agent

语音交互 | 5.3/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5599 字 阅读 →
论文解读

Speech Signals Complement LLMs for Predicting Interpersonal Attraction in Speed Dating

语音交互 | 6.9/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7083 字 阅读 →
论文解读

StanceBench: A Benchmark for Audio LLM-Based Interpersonal Stance Evaluation from Speech

语音属性识别 | 8.6/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8263 字 阅读 →
论文解读

CODA: Cascaded Online Discontinuity-Aware Alignment for Real-Time Image-Based Score Following

音频理解 | 8.1/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6774 字 阅读 →
论文解读

How Meta-Learning Shapes LoRA Adapter Geometry in Speech Deepfake Detection

语音伪造检测 | 5.2/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7098 字 阅读 →
论文解读

MemNMF: Memory-Augmented NMF on LPC Spectra for Anomalous Sound Detection

音频事件检测 | 5.3/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6236 字 阅读 →
论文解读

MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond

语音识别 | 8.2/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6955 字 阅读 →
论文解读

Music-JEPA: Learning a World Model of Sound from Action

音乐转录 | 6.2/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6393 字 阅读 →
论文解读

Phylogenetic signal in marine mammal and bird vocalizations captured by audio foundation models: the limited benefit of domain-specific pretraining

音频理解 | 7.7/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6153 字 阅读 →