论文解读

TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models

音频理解 | 6.1/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8405 字 阅读 →
论文解读

Versatile On-device Adaptation at the Edge by Unifying Few-shot, Zero-shot, Continual, and In-context Learning

元学习 | 8.1/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7787 字 阅读 →
论文解读

Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPO

语音交互 | 7.2/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8276 字 阅读 →
论文解读

Correlation between prosody and pragmatics: A case study of the discourse marker hālā `now' in Persian

语音属性识别 | 4.4/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6591 字 阅读 →
论文解读

Digital Harf: A Clinically Integrated Multimodal AI System for Pervasive Arabic Speech and Language Therapy

语音交互 | 5.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5532 字 阅读 →
论文解读

Enhancing Law-Enforcement Audio Transcription: A LoRA-Based Adaptation of Whisper for BWC Footage

语音识别 | 4.8/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5460 字 阅读 →
论文解读

Integrating Contextual Embeddings into Evaluation of Expressive MIDI Piano Performances

音乐理解 | 7.7/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6828 字 阅读 →
论文解读

RIPPLE: Generating Multi-Channel Phase, Not Recovering It

音频理解 | 7.1/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8301 字 阅读 →
论文解读

SKY-Piano: A Multimodal Piano Performance Dataset

音频理解 | 8.3/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6136 字 阅读 →
论文解读

A Study on Online Mask-based Beamforming Using Per-channel Masking for Spatially Distributed Microphones

语音增强 | 5.7/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4813 字 阅读 →
论文解读

Audio-Anchored Fusion of Multi-Ratio DiT Reconstruction Residuals for Cross-Domain Audio Deepfake Detection

语音伪造检测 | 6.6/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7231 字 阅读 →
论文解读

Detection of AI-generated stems within hybrid human-AI music

CNN | 6.3/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6505 字 阅读 →
论文解读

DuplexGen: Adaptive Synthesis of Human-AI Turn-Taking Dialogues

语音交互 | 6.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5720 字 阅读 →
论文解读

Explicit Note-Event Tokenization and Pitch-Validity Constrained Decoding for MIDI-to-Tablature Transcription

音乐转录 | 6.1/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5320 字 阅读 →
论文解读

Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement

音频分类 | 6.8/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7590 字 阅读 →
论文解读

MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning

音频字幕生成 | 6.3/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6369 字 阅读 →
论文解读

Qwen-Audio-3.0-Gen-Preview Technical Report

音频生成 | 5.6/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7123 字 阅读 →
论文解读

Symphony of Bias: Exploring Gender Associations with Musical Instruments in Multimodal LLMs

音乐理解 | 6.6/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7251 字 阅读 →
论文解读

TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation

音视频生成 | 6.7/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7788 字 阅读 →
论文解读

Unfolded Recursive Expectation-Maximization Neural Network For Speaker Tracking

声源定位 | 5.4/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5982 字 阅读 →