论文解读

APEX: Large-scale Multi-task Aesthetic-Informed Popularity Prediction for AI-Generated Music

音乐理解 | 8.0/10

 · 更新于 2026-10-02 · 约 11 分钟 · 5162 字 阅读 →
论文解读

Assessing the Impact of Noise and Speech Enhancement on the Intelligibility of Speech Codecs

模型评估 | 7.0/10

 · 更新于 2026-10-02 · 约 10 分钟 · 4862 字 阅读 →
论文解读

AsymK-Talker: Real-Time and Long-Horizon Talking Head Generation via Asymmetric Kernel Distillation

语音合成 | 7.5/10

 · 更新于 2026-10-02 · 约 13 分钟 · 6222 字 阅读 →
论文解读

Contrastive Regularization for Accent-Robust ASR

语音识别 | 7.5/10

 · 更新于 2026-10-02 · 约 8 分钟 · 3952 字 阅读 →
论文解读

Cosmodoit: A Python Package for Adaptive, Efficient Pipelining of Feature Extraction from Performed Music

音乐信息检索 | 6.5/10

 · 更新于 2026-10-02 · 约 7 分钟 · 3284 字 阅读 →
论文解读

DECKER: Domain-invariant Embedding for Cross-Keyboard Extraction and Recognition

音频安全 | 7.5/10

 · 更新于 2026-10-02 · 约 12 分钟 · 5862 字 阅读 →
论文解读

Deepfake Audio Detection Using Self-supervised Fusion Representations

音频深度伪造检测 | 7.5/10

 · 更新于 2026-10-02 · 约 9 分钟 · 4110 字 阅读 →
论文解读

Ecologically-Constrained Task Arithmetic for Multi-Taxa Bioacoustic Classifiers Without Shared Data

生物声学 | 8.0/10

 · 更新于 2026-10-02 · 约 13 分钟 · 6468 字 阅读 →
论文解读

Enhancing Self-Supervised Talking Head Forgery Detection via a Training-Free Dual-System Framework

说话头伪造检测 | 7.5/10

 · 更新于 2026-10-02 · 约 14 分钟 · 6781 字 阅读 →
论文解读

Learning Generalizable Action Representations via Pre-training AEMG

生物声学 | 7.5/10

 · 更新于 2026-10-02 · 约 12 分钟 · 5844 字 阅读 →
论文解读

MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model

语音对话系统 | 7.5/10

 · 更新于 2026-10-02 · 约 26 分钟 · 12846 字 阅读 →
论文解读

Mixed-Precision Information Bottlenecks for On-Device Trait-State Disentanglement in Bipolar Agitation Detection

语音生物标志物 | 8.0/10

 · 更新于 2026-10-02 · 约 12 分钟 · 5797 字 阅读 →
论文解读

PHALAR: Phasors for Learned Musical Audio Representations

音乐信息检索 | 8.0/10

 · 更新于 2026-10-02 · 约 12 分钟 · 5654 字 阅读 →
论文解读

Phoneme-Level Deepfake Detection Across Emotional Conditions Using Self-Supervised Embeddings

音频深度伪造检测 | 7.0/10

 · 更新于 2026-10-02 · 约 9 分钟 · 4325 字 阅读 →
论文解读

ReasonAudio: A Benchmark for Evaluating Reasoning Beyond Matching in Text-Audio Retrieval

音频检索 | 7.5/10

 · 更新于 2026-10-02 · 约 11 分钟 · 5435 字 阅读 →
论文解读

Smart Passive Acoustic Monitoring: Embedding a Classifier on AudioMoth Microcontroller

生物声学 | 7.5/10

 · 更新于 2026-10-02 · 约 6 分钟 · 2580 字 阅读 →
论文解读

Stage Light is Sequence$^2$: Multi-Light Control via Imitation Learning

音乐信息检索 | 7.5/10

 · 更新于 2026-10-02 · 约 16 分钟 · 7610 字 阅读 →
论文解读

The TTS-STT Flywheel: Synthetic Entity-Dense Audio Closes the Indic ASR Gap Where Commercial and Open-Source Systems Fail

语音识别 | 8.5/10

 · 更新于 2026-10-02 · 约 13 分钟 · 6496 字 阅读 →
论文解读

Toward Structural Multimodal Representations: Specialization, Selection, and Sparsification via Mixture-of-Experts

多模态模型 | 7.0/10

 · 更新于 2026-10-02 · 约 13 分钟 · 6070 字 阅读 →
论文解读

Towards Open World Sound Event Detection

音频事件检测 | 8.5/10

 · 更新于 2026-10-02 · 约 13 分钟 · 6369 字 阅读 →