论文解读

An Objective Intelligibility Metric Evaluation on Spanish Speech

语音质量评估 | 6.2/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6711 字 阅读 →
论文解读

Anamnesis: An Open-Source Platform for Large-Scale Backstory-Conditioned Survey Simulation

提示学习 | 6.1/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8074 字 阅读 →
论文解读

Anysynth:Zero-Shot Instrument Cloning via In-Context Learning and Asymmetric Hierarchical Guidance

音乐生成 | 6.8/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7974 字 阅读 →
论文解读

ARIMA: Reconstruction-Grounded Predictive Representation Learning for Symbolic Music

自监督学习 | 7.7/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7552 字 阅读 →
论文解读

BackgroundMellow: A Multi-Modal Cohesive Framework for Narrative-Driven Rich Cinematic Soundscape Generation

音频生成 | 7.4/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7705 字 阅读 →
论文解读

BeatEdit: Symbolic Music Generation as Explicit Editing

音乐生成 | 8.9/10

 · 更新于 2026-09-06 · 约 27 分钟 · 13087 字 阅读 →
论文解读

Breaking the Quality--Intelligibility Trade-off in Streaming Target Speaker Extraction via Deep-Feature-Anchored Preference Optimization

语音分离 | 6.3/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7612 字 阅读 →
论文解读

Casting Everything to Online API Services? A Survey of Integrating Localized Speech Recognition Models in Robotic Systems

语音识别 | 5.4/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7257 字 阅读 →
论文解读

CHARM: Charge Calibration and Acoustic Rescue for LLM-based Multimodal Sarcasm Detection

提示学习 | 8.8/10

 · 更新于 2026-09-06 · 约 21 分钟 · 10446 字 阅读 →
论文解读

CoFi-Lite: Pushing the Limits of Ultra-Lightweight Speech Enhancement

语音增强 | 7.3/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6153 字 阅读 →
论文解读

Dance to Music Generation leveraging Pre-training with Unpaired data and Contrastive Alignment

音乐生成 | 7.5/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6661 字 阅读 →
论文解读

Data Augmentation for L2 English Speaking Assessment using TTS

语音质量评估 | 7.0/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7246 字 阅读 →
论文解读

Difference-Driven Gating: Adaptive Feature Fusion for U-Net Decoder

语音分离 | 7.4/10

 · 更新于 2026-09-06 · 约 25 分钟 · 12066 字 阅读 →
论文解读

ECHOv2: Two-Level Band-Splitting Representation Learning for Anomalous Sound Detection

音频事件检测 | 8.2/10

 · 更新于 2026-09-06 · 约 18 分钟 · 8793 字 阅读 →
论文解读

Efficiently Adapting Spoken Language Models for the Singaporean Context

语音交互 | 6.5/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8212 字 阅读 →
论文解读

Encoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language Models

音频理解 | 6.4/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7564 字 阅读 →
论文解读

Evaluating SSL and ViViT Architectures for Cross-Corpus Audio MOS Prediction via LODO Validation

语音质量评估 | 8.3/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7355 字 阅读 →
论文解读

Evidence Subspace Projection: Measuring How Much Evidence Explains Deepfake Detection in Self-Supervised Speech Models

语音伪造检测 | 8.1/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8254 字 阅读 →
论文解读

FdAudio: MeanFlow-Anchored Fréchet-Distance Post-Training for One-Step Text-to-Audio Generation

音频生成 | 8.6/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8443 字 阅读 →
论文解读

GigaAM Multilingual: Foundation Model for Underrepresented Languages

语音识别 | 8.1/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6967 字 阅读 →