论文解读

Advancing Electrolaryngeal Speech Enhancement Through Speech-Text Representation Learning

语音增强 | 7.1/10

 · 更新于 2026-09-09 · 约 17 分钟 · 8092 字 阅读 →
论文解读

AI Slop or AI-enhancement? Student perceptions of AI-generated media for an English for Academic Purposes course

AI Slop or AI-enhancement? Student perceptions of AI-generated media for an English for Academic Purposes course

 · 更新于 2026-09-09 · 约 13 分钟 · 6238 字 阅读 →
论文解读

AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling

多模态模型 | 7/10

 · 更新于 2026-09-09 · 约 15 分钟 · 7267 字 阅读 →
论文解读

Beyond the Mouth: Upper-Face Affective Cues in Audiovisual Sentence Recognition under Acoustic Uncertainty

语音识别 | 5.5/10

 · 更新于 2026-09-09 · 约 10 分钟 · 4920 字 阅读 →
论文解读

Context-aware child-directed speech detection from long-form recordings

自监督学习 | 8.5/10

 · 更新于 2026-09-09 · 约 11 分钟 · 5403 字 阅读 →
论文解读

DAStatFormer: A Hybrid Multibranch Transformer with Statistical Feature Integration for DAS-Based Pattern Recognitions

音频事件检测 | 6.4/10

 · 更新于 2026-09-09 · 约 10 分钟 · 4723 字 阅读 →
论文解读

Description and Discussion on DCASE 2026 Challenge Task 2: Noise-aware Unsupervised Anomalous Sound Detection for Machine Condition Monitoring

无监督学习 | 7.2/10

 · 更新于 2026-09-09 · 约 10 分钟 · 4656 字 阅读 →
论文解读

DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech

语音合成 | 7.1/10

 · 更新于 2026-09-09 · 约 14 分钟 · 6680 字 阅读 →
论文解读

Dynamic Interaction-Aware and Causality-Disentangled Framework for Multimodal Sentiment Analysis

多模态模型 | 7.8/10

 · 更新于 2026-09-09 · 约 12 分钟 · 6012 字 阅读 →
论文解读

Echo: A Joint-Embedding Predictive Architecture for Speaker Diarization and Speech Recognition in a Shared Latent Space

语音识别 | 7/10

 · 更新于 2026-09-09 · 约 15 分钟 · 7107 字 阅读 →
论文解读

HAIM: Human-AI Music Datasets for AI Music Production Tracking Benchmark

HAIM: Human-AI Music Datasets for AI Music Production Tracking Benchmark

 · 更新于 2026-09-09 · 约 10 分钟 · 4653 字 阅读 →
论文解读

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions

音乐生成 | 7.3/10

 · 更新于 2026-09-09 · 约 13 分钟 · 6236 字 阅读 →
论文解读

Kinship Verification Using Voice

声纹识别 | 6.9/10

 · 更新于 2026-09-09 · 约 11 分钟 · 5448 字 阅读 →
论文解读

Local Diagnostics of Continuous Normalizing Flow for Out-of-Distribution Detection

语音合成 | 8.1/10

 · 更新于 2026-09-09 · 约 12 分钟 · 5709 字 阅读 →
论文解读

MelT: GEMM-Native NDFT for Efficient Single-Stage Audio Frontends on Modern Accelerators

信号处理基础 | 7.3/10

 · 更新于 2026-09-09 · 约 12 分钟 · 5889 字 阅读 →
论文解读

MOSS-Audio Technical Report

语音识别 | 9.2/10

 · 更新于 2026-09-09 · 约 12 分钟 · 5940 字 阅读 →
论文解读

Multimodal Music Recommendation System using LLMs

音乐推荐 | 10/10

 · 更新于 2026-09-09 · 约 15 分钟 · 7397 字 阅读 →
论文解读

MURMUR: An Efficient Inference System for Long-Form ASR

语音识别 | 8.3/10

 · 更新于 2026-09-09 · 约 4 分钟 · 1773 字 阅读 →
论文解读

Parameter-efficient Dual-encoder Architecture with Differentiable Choquet Integral Fusion for Underwater Acoustic Classification

音频分类 | 6.4/10

 · 更新于 2026-09-09 · 约 12 分钟 · 5654 字 阅读 →
论文解读

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects

语音识别 | 8.8/10

 · 更新于 2026-09-09 · 约 12 分钟 · 5668 字 阅读 →