论文解读

Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories

Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories

 · 更新于 2026-09-09 · 约 13 分钟 · 6279 字 阅读 →
论文解读

Deploying Speech-Driven 3D Facial Animation in Unreal Engine for Production-Ready Digital Humans

语音合成 | 6.6/10

 · 更新于 2026-09-09 · 约 24 分钟 · 11714 字 阅读 →
论文解读

DeRA-MOS: Optimizing Text-to-Music Evaluation via Decoupled Listwise Ranking and Modality Alignment

音乐评估 | 8.2/10

 · 更新于 2026-09-09 · 约 11 分钟 · 5161 字 阅读 →
论文解读

Dual-Branch Gated Fusion for Open-Set Audio Deepfake Source Tracing

音频深度伪造检测 | 7.8/10

 · 更新于 2026-09-09 · 约 13 分钟 · 6305 字 阅读 →
论文解读

Enhancing Multilingual LLM-based ASR with Mixture of Experts and Dynamic Downsampling

Enhancing Multilingual LLM-based ASR with Mixture of Experts and Dynamic Downsampling

 · 更新于 2026-09-09 · 约 21 分钟 · 10095 字 阅读 →
论文解读

Entropy-Aware Domain-Routed Mixture-of-Experts Speech-LLM Framework: A Case Study of Multi-Domain Child-Adult ASR

语音识别 | 7.5/10

 · 更新于 2026-09-09 · 约 13 分钟 · 6225 字 阅读 →
论文解读

Ethical and Technical Limits of Deepfake Speech Datasets

Ethical and Technical Limits of Deepfake Speech Datasets

 · 更新于 2026-09-09 · 约 23 分钟 · 11092 字 阅读 →
论文解读

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs

语音识别 | 6.5/10

 · 更新于 2026-09-09 · 约 14 分钟 · 6823 字 阅读 →
论文解读

GC-LoRA: Gated Convolutional LoRA for Parameter-Efficient Acoustic Adaptation

语音识别 | 7.6/10

 · 更新于 2026-09-09 · 约 9 分钟 · 4079 字 阅读 →
论文解读

GlobeAudio: A Multilingual Multicultural Benchmark for Naturalistic Evaluation of Large Audio-Language Models

语音识别 | 7.9/10

 · 更新于 2026-09-09 · 约 11 分钟 · 5486 字 阅读 →
论文解读

Inside the Latent Flow: Causal Deciphering of Attention Dynamics in Audio Separation Foundation Models

语音分离 | 7.3/10

 · 更新于 2026-09-09 · 约 14 分钟 · 6713 字 阅读 →
论文解读

KFC-KWS: Keyframe Fusion with CTC for User-Defined Keyword Spotting

关键词检测 | 7.6/10

 · 更新于 2026-09-09 · 约 11 分钟 · 5345 字 阅读 →
论文解读

Linguistically Augmented Audio Speech Data (LinguAS)

语音伪造检测 | 7.5/10

 · 更新于 2026-09-09 · 约 10 分钟 · 4691 字 阅读 →
论文解读

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling

语音识别 | 8.6/10

 · 更新于 2026-09-09 · 约 12 分钟 · 5701 字 阅读 →
论文解读

Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models

语音对话系统 | 9.3/10

 · 更新于 2026-09-09 · 约 18 分钟 · 8738 字 阅读 →
论文解读

Multilingual Word-Level Forced Alignment with Self-Supervised Representations and Learned Dynamic Programming

自监督学习 | 6.3/10

 · 更新于 2026-09-09 · 约 11 分钟 · 5292 字 阅读 →
论文解读

OmniCap-IF: Benchmarking and Improving Instruction Following Abilities for Omni-Video Captioning

语音生成 | 9.1/10

 · 更新于 2026-09-09 · 约 13 分钟 · 6024 字 阅读 →
论文解读

Optimality of FSQ Tokens for Continuous Diffusion for Categorical Data with Application to Text-to-Speech

语音合成 | 7.3/10

 · 更新于 2026-09-09 · 约 11 分钟 · 5034 字 阅读 →
论文解读

Optimizing 2D Input Representations and Sub-phase Fusion Strategies for Differential Diagnosis of Asthma and COPD Using CNN- and GRU-Based Networks

数据增强 | 6.8/10

 · 更新于 2026-09-09 · 约 48 分钟 · 23710 字 阅读 →
论文解读

Overview of ESDD2: Environment-Aware Speech and Sound Deepfake Detection Challenge

数据增强 | 6.3/10

 · 更新于 2026-09-09 · 约 18 分钟 · 8529 字 阅读 →