论文解读

BackgroundMellow: A Multi-Modal Cohesive Framework for Narrative-Driven Rich Cinematic Soundscape Generation

音频生成 | 7.4/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7705 字 阅读 →
论文解读

BeatEdit: Symbolic Music Generation as Explicit Editing

音乐生成 | 8.9/10

 · 更新于 2026-09-25 · 约 27 分钟 · 13087 字 阅读 →
论文解读

Breaking the Quality--Intelligibility Trade-off in Streaming Target Speaker Extraction via Deep-Feature-Anchored Preference Optimization

语音分离 | 6.3/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7612 字 阅读 →
论文解读

Casting Everything to Online API Services? A Survey of Integrating Localized Speech Recognition Models in Robotic Systems

语音识别 | 5.4/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7257 字 阅读 →
论文解读

Data Augmentation for L2 English Speaking Assessment using TTS

语音质量评估 | 7.0/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7246 字 阅读 →
论文解读

Difference-Driven Gating: Adaptive Feature Fusion for U-Net Decoder

语音分离 | 7.4/10

 · 更新于 2026-09-25 · 约 25 分钟 · 12066 字 阅读 →
论文解读

ECHOv2: Two-Level Band-Splitting Representation Learning for Anomalous Sound Detection

音频事件检测 | 8.2/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8793 字 阅读 →
论文解读

Efficiently Adapting Spoken Language Models for the Singaporean Context

语音交互 | 6.5/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8212 字 阅读 →
论文解读

Encoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language Models

音频理解 | 6.4/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7564 字 阅读 →
论文解读

Evaluating SSL and ViViT Architectures for Cross-Corpus Audio MOS Prediction via LODO Validation

语音质量评估 | 8.3/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7355 字 阅读 →
论文解读

Evidence Subspace Projection: Measuring How Much Evidence Explains Deepfake Detection in Self-Supervised Speech Models

语音伪造检测 | 8.1/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8254 字 阅读 →
论文解读

GigaChat Audio: Time-aware Large Audio Language Model

音频理解 | 7.4/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6899 字 阅读 →
论文解读

Graph Representation of RaagBase: A Unique Dataset for Hindustani Music

音乐理解 | 5.7/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5708 字 阅读 →
论文解读

Learn2Chat: Rethinking Dyadic Talking Heads via Interaction-Modulated Monologic Priors

音视频生成 | 6.8/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6464 字 阅读 →
论文解读

LightMem-Ego: Your AI Memory for Everyday Life

流式处理 | 5.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5925 字 阅读 →
论文解读

Listen to the Features: Voice Anonymization Driven by Content Embedding Matching over Signal Reconstruction

语音克隆 | 6.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6792 字 阅读 →
论文解读

Local Multimodal Music Alignment from Global Supervision

对比学习 | 7.6/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7732 字 阅读 →
论文解读

LOGOS: A Living Logic for AI Agent Teams That Evolve With Humans

多模态模型 | 6.1/10

 · 更新于 2026-09-25 · 约 21 分钟 · 10138 字 阅读 →
论文解读

MRUF: Multi-granularity Routing with Uncertainty-Aware Fusion for Robust Multimodal Sentiment Analysis

多模态模型 | 5.9/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8734 字 阅读 →
论文解读

MusicMark: A Robust Generative Watermarking Framework for Music Generation

音频水印 | 7.3/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8863 字 阅读 →