论文解读

AudioMap: Cloze-and-Choice Reinforcement Learning for Time-Aware Dense Audio Captioning

音频字幕生成 | 7.9/10

 · 更新于 2026-09-25 · 约 5 分钟 · 2475 字 阅读 →
论文解读

Beyond Piano: Cross-Instrument MIDI Velocity Estimation via Differentiable SoundFont Proxies

音乐理解 | 7.3/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8783 字 阅读 →
论文解读

Dynamic Clustering for Cross-Segment Permutation Alignment in Long Speech Separation

语音分离 | 6.2/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6452 字 阅读 →
论文解读

From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs

音频理解 | 7.4/10

 · 更新于 2026-09-25 · 约 5 分钟 · 2056 字 阅读 →
论文解读

Investigating Multimodal Informativity under Different Partner Visibility Conditions in Video-Mediated Dialogue

多模态模型 | 5.9/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7905 字 阅读 →
论文解读

MusicLayout: Explicit Structural Planning for Controllable Text-to-Music Generation

音乐生成 | 6.3/10

 · 更新于 2026-09-25 · 约 24 分钟 · 11885 字 阅读 →
论文解读

Neural Array-Generic Direction-of-Arrival Estimation Exploiting Array Transfer Functions

声源定位 | 6.4/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6231 字 阅读 →
论文解读

SAMOT: State-Aware Step Modulation and Optimal Transport Matching for Audio-Visual Instance Segmentation

音视频理解 | 7.7/10

 · 更新于 2026-09-25 · 约 25 分钟 · 12060 字 阅读 →
论文解读

The Voiceprint Fallacy: Why Voices Are Not Unique Biometric Imprints

说话人验证 | 6.0/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6330 字 阅读 →
论文解读

From Prompting to Describing: A Cross-Cultural Study of Language for AI-Generated Music

音乐理解 | 5.5/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7605 字 阅读 →
论文解读

How Much AI Is in This Track? Quantifying the Proportion of AI-Generated Stems in Hybrid Music Mixtures

音频伪造检测 | 8.3/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6888 字 阅读 →
论文解读

MI-MIDI: Mechanistic Interpretability of Text-to-MIDI Generation Models via Probing, Lenses and Steering

音乐生成 | 7.1/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6408 字 阅读 →
论文解读

MMAG: A Multi-Control Mixed Audio Generation Benchmark

音频生成 | 6.8/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7097 字 阅读 →
论文解读

Separating Decision-Rule Misalignment from Readout-Coverage Limitations in Speech Language Models

语音情感识别 | 6.8/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7558 字 阅读 →
论文解读

Beyond Residual Connections: Manifold-Constrained Hyper-Connections for Robust Speaker Representation Learning

说话人验证 | 6.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4473 字 阅读 →
论文解读

Bias Analysis of L2 Speaking Assessment Systems Using Concept Activation Vectors

语音质量评估 | 6.7/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7166 字 阅读 →
论文解读

Diff2Mix: Controllable Music Mixing via Diffusion Models and Differentiable Audio Effects

音频生成 | 7.5/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7204 字 阅读 →
论文解读

Explicit and Stable Pseudospectral Time-Domain Method for the Föppl-von Kármán Equations

音频理解 | 7.2/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6256 字 阅读 →
论文解读

LILAC: An Idempotent Neural Speech Codec

语音编码 | 8.1/10

 · 更新于 2026-09-25 · 约 12 分钟 · 6008 字 阅读 →
论文解读

The interface of intonation and lexical tone: Boundary phenomena in Mandarin varieties

语音属性识别 | 5.6/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4830 字 阅读 →