论文解读

SceneBind: Binding What and Where Across Vision, Audio and Language

音视频理解 | 6.6/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6403 字 阅读 →
论文解读

Tight-Frame Reconstruction for Acoustic Intensity Estimation Using Cardioid Microphone Pairs

声源定位 | 6.8/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7054 字 阅读 →
论文解读

Dual-BEATs: Unlocking Zero-Shot Stereo Audio Perception in Audio Large Language Models via Dithering

多模态模型 | 7.1/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7057 字 阅读 →
论文解读

Learning-based Physics-Constrained Neural Kernel for Sound Field Estimation With Source-Position-Dependent Directional Weighting

声源定位 | 5.2/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6054 字 阅读 →
论文解读

INFER: Learning Implicit Neural Frequency Response Fields for Confined Acoustic Environments

空间音频 | 6.4/10

 · 更新于 2026-09-24 · 约 6 分钟 · 2761 字 阅读 →
论文解读

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments

声源定位 | 8.1/10

 · 更新于 2026-09-24 · 约 17 分钟 · 8152 字 阅读 →
论文解读

PhaseCoder: Microphone Geometry-Agnostic Spatial Audio Understanding for Multimodal LLMs

空间音频 | 8.7/10

 · 更新于 2026-09-24 · 约 17 分钟 · 8299 字 阅读 →
论文解读

Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer

音视频生成 | 6.6/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6714 字 阅读 →
论文解读

Spatial Speech Perception Systems: A Survey of Sound Source Localization, Directional Enhancement, and Speech Recognition

声源定位 | 4.1/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5837 字 阅读 →
论文解读

Evaluation of Head-Related Transfer Functions Across Five Levels of Individualisation in Virtual Reality

空间音频 | 7.6/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5571 字 阅读 →
论文解读

Perceptual Evaluation of Higher-Order Ambisonic Codecs on Both Synthetic Mixing and Native Recordings

音频编码 | 8/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6403 字 阅读 →
论文解读

Generalised Transcoding Framework for Arbitrary Spatial Audio Capture and Playback Formats

空间音频 | 8.7/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5840 字 阅读 →
论文解读

Spatio-Temporal Audio Language Modeling for Dynamic Sound Sources

音频问答 | 6.9/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5968 字 阅读 →
论文解读

Sensitivity Analysis of Generative Spatial Audio Metrics: A Study on Responsiveness, Smoothness, and Symmetry

音频生成 | 7.2/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5919 字 阅读 →
论文解读

Flow-HOA: Generative Joint Optimization for Ambisonics Encoding via Flow Matching

空间音频 | 7.9/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6364 字 阅读 →
论文解读

SHB-AE: Spherical harmonic beamforming based Ambisonics encoding and upscaling method for smartphone microphone array

音频编码 | 6.7/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5593 字 阅读 →
论文解读

From Numbers to Perception, Energy Decay Curves Prediction

空间音频 | 7.2/10

 · 更新于 2026-09-24 · 约 17 分钟 · 8331 字 阅读 →
论文解读

Spatial Power Estimation via Riemannian Covariance Matching

声源定位 | 6.5/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7380 字 阅读 →
论文解读

NDF+: Joint Neural Directional Filtering and Diffuse Sound Extraction

空间音频 | 6.5/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5284 字 阅读 →
会议任务专题

ICLR 2026 - 空间音频

共 1 篇 ICLR 2026 空间音频 方向论文

 · 更新于 2026-09-24 · 约 4 分钟 · 1689 字 阅读 →