论文解读

All That Glitters Is Not Audio: Rethinking Text Priors and Audio Reliance in Audio-Language Evaluation

音频问答 | 6.5/10

 · 更新于 2026-09-10 · 约 9 分钟 · 4036 字 阅读 →
论文解读

An event-based sequence modeling approach to recognizing non-triad chords with oversegmentation minimization

音乐理解 | 7.5/10

 · 更新于 2026-09-10 · 约 9 分钟 · 4400 字 阅读 →
论文解读

CineAGI: Character-Consistent Movie Creation through LLM-Orchestrated Multi-Modal Generation and Cross-Scene Integration

跨模态 | 8.0/10

 · 更新于 2026-09-10 · 约 9 分钟 · 4377 字 阅读 →
论文解读

Come Together: Analyzing Popular Songs Through Statistical Embeddings

音乐信息检索 | 6.5/10

 · 更新于 2026-09-10 · 约 8 分钟 · 3526 字 阅读 →
论文解读

Comparison of sEMG Encoding Accuracy Across Speech Modes Using Articulatory and Phoneme Features

语音生物标志物 | 8.0/10

 · 更新于 2026-09-10 · 约 9 分钟 · 4365 字 阅读 →
论文解读

Explainable AI in Speaker Recognition -- Making Latent Representations Understandable

说话人识别 | 7.5/10

 · 更新于 2026-09-10 · 约 8 分钟 · 3672 字 阅读 →
论文解读

Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation

音视频 | 8.5/10

 · 更新于 2026-09-10 · 约 12 分钟 · 5626 字 阅读 →
论文解读

HeadRouter: Dynamic Head-Weight Routing for Task-Adaptive Audio Token Pruning in Large Audio Language Models

音频大模型 | 8.0/10

 · 更新于 2026-09-10 · 约 11 分钟 · 5044 字 阅读 →
论文解读

Latent-Hysteresis Graph ODEs: Modeling Coupled Topology-Feature Evolution via Continuous Phase Transitions

图神经网络 | 8.0/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4777 字 阅读 →
论文解读

Listening with Time: Precise Temporal Awareness for Long-Form Audio Understanding

音频场景理解 | 8.0/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4909 字 阅读 →
论文解读

MAGIC-TTS: Fine-Grained Controllable Speech Synthesis with Explicit Local Duration and Pause Control

语音合成 | 7.0/10

 · 更新于 2026-09-10 · 约 9 分钟 · 4497 字 阅读 →
论文解读

Meta-Ensemble Learning with Diverse Data Splits for Improved Respiratory Sound Classification

音频分类 | 8.0/10

 · 更新于 2026-09-10 · 约 9 分钟 · 4218 字 阅读 →
论文解读

Opening the Design Space: Two Years of Performance with Intelligent Musical Instruments

音乐生成 | 6.5/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4910 字 阅读 →
论文解读

Predictive Directional Selective Fixed-Filter Active Noise Control for Moving Sources via a Convolutional Recurrent Neural Network

声源定位 | 7.5/10

 · 更新于 2026-09-10 · 约 8 分钟 · 3941 字 阅读 →
论文解读

Psychologically-Grounded Graph Modeling for Interpretable Depression Detection

语音情感识别 | 8.0/10

 · 更新于 2026-09-10 · 约 18 分钟 · 8689 字 阅读 →
论文解读

RAS: a Reliability Oriented Metric for Automatic Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-10 · 约 9 分钟 · 4410 字 阅读 →
论文解读

Robust Audio-Text Retrieval via Cross-Modal Attention and Hybrid Loss

音频检索 | 7.5/10

 · 更新于 2026-09-10 · 约 9 分钟 · 4499 字 阅读 →
论文解读

RTCFake: Speech Deepfake Detection in Real-Time Communication

语音伪造检测 | 7.0/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4522 字 阅读 →
论文解读

Scaling Properties of Continuous Diffusion Spoken Language Models

语音生成 | 8.0/10

 · 更新于 2026-09-10 · 约 13 分钟 · 6298 字 阅读 →
论文解读

Spectro-Temporal Modulation Representation Framework for Human-Imitated Speech Detection

语音伪造检测 | 6.5/10

 · 更新于 2026-09-10 · 约 7 分钟 · 3399 字 阅读 →