论文解读

Bridging Piano Transcription and Rendering via Disentangled Score Content and Style

音乐信息检索 | 8.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5197 字 阅读 →
论文解读

Human Behavior Atlas: Benchmarking Unified Psychological And Social Behavior Understanding

多模态模型 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5900 字 阅读 →
论文解读

Speech World Model: Causal State–Action Planning with Explicit Reasoning for Speech

语音情感识别 语音对话系统 | 9.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5094 字 阅读 →
论文解读

SpeechOp: Inference-Time Task Composition for Generative Speech Processing

语音增强 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5070 字 阅读 →
论文解读

WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM

音频检索 视频检索 | 8.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5230 字 阅读 →
论文解读

EmoTransCap: Dataset and Pipeline for Emotion Transition-Aware Speech Captioning in Discourses

语音情感识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5250 字 阅读 →
论文解读

A Task-Aware Dual-Level Self-Supervised Learning Method for Effective Sound Event Detection

音频事件检测 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4113 字 阅读 →
论文解读

ACAVCaps: Enabling Large-Scale Training for Fine-Grained and Diverse Audio Understanding

音频分类 | 8.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3768 字 阅读 →
论文解读

AccLID: Accent-aware Language Identification for Robust Multilingual Speech Recognition

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4432 字 阅读 →
论文解读

Advanced modeling of interlanguage speech intelligibility benefit with L1-L2 multi-task learning using differentiable K-means for accent-robust discrete token-based ASR

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4144 字 阅读 →
论文解读

An Envelope Separation Aided Multi-Task Learning Model for Blind Source Counting and Localization

声源定位 | 6.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4186 字 阅读 →
论文解读

Assessing the Impact of Speaker Identity in Speech Spoofing Detection

音频深度伪造检测 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4764 字 阅读 →
论文解读

ATOM: Adaptive Token-Level Optimal Transport Mixup for Speech Translation

语音翻译 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4810 字 阅读 →
论文解读

Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding

语音编码器 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4347 字 阅读 →
论文解读

Audio-Visual Feature Fusion for Calibrating Relevance Scores of Video Moment Retrieval

视频片段检索 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5439 字 阅读 →
论文解读

Auxiliary Multi-Label Training For Improving the Robustness of Audio Deepfake Detection on AI-Processed Data

音频深度伪造检测 | 6.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3885 字 阅读 →
论文解读

Beyond Global Emotion: Fine-Grained Emotional Speech Synthesis with Dynamic Word-Level Modulation

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4555 字 阅读 →
论文解读

Brainprint-Modulated Target Speaker Extraction

语音分离 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4574 字 阅读 →
论文解读

CALM: Joint Contextual Acoustic-Linguistic Modeling for Personalization of Multi-Speaker ASR

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5139 字 阅读 →
论文解读

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources

音频场景理解 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4177 字 阅读 →