论文解读

Context-aware child-directed speech detection from long-form recordings

自监督学习 | 8.5/10

 · 更新于 2026-09-30 · 约 11 分钟 · 5403 字 阅读 →
论文解读

DAStatFormer: A Hybrid Multibranch Transformer with Statistical Feature Integration for DAS-Based Pattern Recognitions

音频事件检测 | 6.4/10

 · 更新于 2026-09-30 · 约 10 分钟 · 4723 字 阅读 →
论文解读

Description and Discussion on DCASE 2026 Challenge Task 2: Noise-aware Unsupervised Anomalous Sound Detection for Machine Condition Monitoring

无监督学习 | 7.2/10

 · 更新于 2026-09-30 · 约 10 分钟 · 4656 字 阅读 →
论文解读

DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech

语音合成 | 7.1/10

 · 更新于 2026-09-30 · 约 14 分钟 · 6680 字 阅读 →
论文解读

Dynamic Interaction-Aware and Causality-Disentangled Framework for Multimodal Sentiment Analysis

多模态模型 | 7.8/10

 · 更新于 2026-09-30 · 约 12 分钟 · 6012 字 阅读 →
论文解读

Echo: A Joint-Embedding Predictive Architecture for Speaker Diarization and Speech Recognition in a Shared Latent Space

语音识别 | 7/10

 · 更新于 2026-09-30 · 约 15 分钟 · 7107 字 阅读 →
论文解读

HAIM: Human-AI Music Datasets for AI Music Production Tracking Benchmark

HAIM: Human-AI Music Datasets for AI Music Production Tracking Benchmark

 · 更新于 2026-09-30 · 约 10 分钟 · 4653 字 阅读 →
论文解读

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions

音乐生成 | 7.3/10

 · 更新于 2026-09-30 · 约 13 分钟 · 6236 字 阅读 →
论文解读

Kinship Verification Using Voice

声纹识别 | 6.9/10

 · 更新于 2026-09-30 · 约 11 分钟 · 5448 字 阅读 →
论文解读

Local Diagnostics of Continuous Normalizing Flow for Out-of-Distribution Detection

语音合成 | 8.1/10

 · 更新于 2026-09-30 · 约 12 分钟 · 5709 字 阅读 →
论文解读

MelT: GEMM-Native NDFT for Efficient Single-Stage Audio Frontends on Modern Accelerators

信号处理基础 | 7.3/10

 · 更新于 2026-09-30 · 约 12 分钟 · 5889 字 阅读 →
论文解读

MOSS-Audio Technical Report

语音识别 | 9.2/10

 · 更新于 2026-09-30 · 约 12 分钟 · 5940 字 阅读 →
论文解读

Multimodal Music Recommendation System using LLMs

音乐推荐 | 10/10

 · 更新于 2026-09-30 · 约 15 分钟 · 7397 字 阅读 →
论文解读

MURMUR: An Efficient Inference System for Long-Form ASR

语音识别 | 8.3/10

 · 更新于 2026-09-30 · 约 4 分钟 · 1773 字 阅读 →
论文解读

Parameter-efficient Dual-encoder Architecture with Differentiable Choquet Integral Fusion for Underwater Acoustic Classification

音频分类 | 6.4/10

 · 更新于 2026-09-30 · 约 12 分钟 · 5654 字 阅读 →
论文解读

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects

语音识别 | 8.8/10

 · 更新于 2026-09-30 · 约 12 分钟 · 5668 字 阅读 →
论文解读

Privacy-preserving Prosody Representation Learning

自监督学习 | 4.9/10

 · 更新于 2026-09-30 · 约 11 分钟 · 5261 字 阅读 →
论文解读

Project SPARROW and the Future of Conservation Technology

计算机视觉 | 10/10

 · 更新于 2026-09-30 · 约 13 分钟 · 6098 字 阅读 →
论文解读

Quality Audio Prototyping: a prototype system for unified sound retrieval and procedural generation

音频检索 | 6.9/10

 · 更新于 2026-09-30 · 约 10 分钟 · 4629 字 阅读 →
论文解读

RRP-Voice: A Longitudinal Dataset and Benchmark for Recurrent Respiratory Papillomatosis Detection

数据集 | 8.3/10

 · 更新于 2026-09-30 · 约 17 分钟 · 8064 字 阅读 →