论文解读

Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis

音频交互 | 7.0/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6613 字 阅读 →
论文解读

Qwen-Audio-3.0-Gen-Preview Technical Report

音频生成 | 5.6/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7123 字 阅读 →
论文解读

Symphony of Bias: Exploring Gender Associations with Musical Instruments in Multimodal LLMs

音乐理解 | 6.6/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7251 字 阅读 →
论文解读

TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation

音视频生成 | 6.7/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7788 字 阅读 →
论文解读

ThinkOmni: A Reasoning-Driven Omni-Modal LLM Framework for Audio Forgery Detection and Localization

音频伪造检测 | 8.3/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7713 字 阅读 →
论文解读

Unfolded Recursive Expectation-Maximization Neural Network For Speaker Tracking

声源定位 | 5.4/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5982 字 阅读 →
论文解读

Voice Memory for Agentic Speech Recognition

语音识别 | 8.2/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8134 字 阅读 →
论文解读

A Cross-lingual Comparison of Human and Classification Model Entrainment Behavior in Code-switched Speech Settings

Transformer | 5.1/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7784 字 阅读 →
论文解读

AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition

语音情感识别 | 5.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5579 字 阅读 →
论文解读

AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities

多模态模型 | 8.3/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5386 字 阅读 →
论文解读

Depression Markers in Speech: An Approach based on Tract Variables Dynamics

语音属性识别 | 5.1/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6197 字 阅读 →
论文解读

Device Invariance using Domain Adaptation on Acoustic Scene Classification

领域适应 | 4.2/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6363 字 阅读 →
论文解读

DynaBridge: Dynamic Summary-Guided Cross-Task Multimodal Fusion for DASS-Structured Mental Health Assessment

音视频理解 | 4.9/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8175 字 阅读 →
论文解读

Evaluation of forced alignment of code-mixed speech: the case of Hindi-English

语音识别 | 5.4/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4432 字 阅读 →
论文解读

Extracting Voice Styles from Frozen TTS Models via Gradient-Based Inverse Optimization

语音克隆 | 7.9/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5923 字 阅读 →
论文解读

Finding the noise: Zero-shot AI Music Detection

无监督学习 | 5.1/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6037 字 阅读 →
论文解读

From Semantics to Readout: Mechanistic Understanding of Audio Tokens after Fine-Tuning for Temporal Audio Grounding

音频理解 | 5.1/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6416 字 阅读 →
论文解读

GraphIDyOM: A graph-native Python reimplementation of IDyOM for musical expectation modelling

音乐理解 | 7.8/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5997 字 阅读 →
论文解读

Joint Text-Audio Alignment for EEG-to-Text Decoding in Chinese Speech Production and Perception

语音交互 | 7.0/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7535 字 阅读 →
论文解读

Multi-Phonation Graph Learning with Self-Supervised Speech Embeddings for ALS Detection and Progression Prediction

语音属性识别 | 5.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6989 字 阅读 →