论文解读

Kinetic-Optimal Scheduling with Moment Correction for Metric-Induced Discrete Flow Matching in Zero-Shot Text-to-Speech

语音合成 | 5.5/10

 · 更新于 2026-09-10 · 约 20 分钟 · 9695 字 阅读 →
论文解读

Latent Secret Spin: Keyed Orthogonal Rotations for Blind Speech Watermarking in Anisotropic Latent Spaces

音频水印 | 5.5/10

 · 更新于 2026-09-10 · 约 16 分钟 · 7659 字 阅读 →
论文解读

Low-Cost Detection of Degraded Voice Clones via Source-Output Acoustic Consistency

语音伪造检测 | 5.3/10

 · 更新于 2026-09-10 · 约 16 分钟 · 7610 字 阅读 →
论文解读

Mitigating Multimodal Inconsistency via Cognitive Dual-Pathway Reasoning for Intent Recognition

意图识别 | 7.0/10

 · 更新于 2026-09-10 · 约 19 分钟 · 9463 字 阅读 →
论文解读

Multi-layer attentive probing improves transfer of audio representations for bioacoustics

生物声学 音频分类 | 4.0/10

 · 更新于 2026-09-10 · 约 16 分钟 · 7879 字 阅读 →
论文解读

Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search

基准测试 | 6.0/10

 · 更新于 2026-09-10 · 约 16 分钟 · 7890 字 阅读 →
论文解读

Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization

基准测试 | 6.5/10

 · 更新于 2026-09-10 · 约 18 分钟 · 8876 字 阅读 →
论文解读

Online Segmented Beamforming via Dynamic Programming

声源定位 | 6.0/10

 · 更新于 2026-09-10 · 约 14 分钟 · 6970 字 阅读 →
论文解读

PoDAR: Power-Disentangled Audio Representation for Generative Modeling

语音合成 | 7.3/10

 · 更新于 2026-09-10 · 约 18 分钟 · 8991 字 阅读 →
论文解读

Polyphonia: Zero-Shot Timbre Transfer in Polyphonic Music with Acoustic-Informed Attention Calibration

音乐生成 | 7.5/10

 · 更新于 2026-09-10 · 约 18 分钟 · 8830 字 阅读 →
论文解读

Probing Cross-modal Information Hubs in Audio-Visual LLMs

模型分析 | 6.5/10

 · 更新于 2026-09-10 · 约 19 分钟 · 9363 字 阅读 →
论文解读

RADAR Challenge 2026: Robust Audio Deepfake Recognition under Media Transformations

音频深度伪造检测 | 6.0/10

 · 更新于 2026-09-10 · 约 16 分钟 · 7522 字 阅读 →
论文解读

Reducing Linguistic Hallucination in LM-Based Speech Enhancement via Noise-Invariant Acoustic-Semantic Distillation

语音增强 | 7.2/10

 · 更新于 2026-09-10 · 约 21 分钟 · 10260 字 阅读 →
论文解读

Remix the Timbre: Diffusion-Based Style Transfer Across Polyphonic Stems

音色迁移 | 5.5/10

 · 更新于 2026-09-10 · 约 18 分钟 · 8892 字 阅读 →
论文解读

Responsible Benchmarking of Fairness for Automatic Speech Recognition

语音识别 | 5.0/10

 · 更新于 2026-09-10 · 约 16 分钟 · 7970 字 阅读 →
论文解读

Rethinking Entropy Minimization in Test-Time Adaptation for Autoregressive Models

语音识别 | 6.0/10

 · 更新于 2026-09-10 · 约 19 分钟 · 9094 字 阅读 →
论文解读

Separate First, Fuse Later: Mitigating Cross-Modal Interference in Audio-Visual LLMs Reasoning with Modality-Specific Chain-of-Thought

音视频问答 | 6.0/10

 · 更新于 2026-09-10 · 约 20 分钟 · 9627 字 阅读 →
论文解读

SF-Flow: Sound field magnitude estimation via flow matching guided by sparse measurements

空间音频 | 6.8/10

 · 更新于 2026-09-10 · 约 16 分钟 · 7898 字 阅读 →
论文解读

ShipEcho -- An Interactive Tool for Global Mapping of Underwater Radiated Noise from Vessels

水下声学 | 6.0/10

 · 更新于 2026-09-10 · 约 15 分钟 · 7190 字 阅读 →
论文解读

Single-Microphone Audio Point Source Discriminative Localization From Reverberation Late Tail Estimation

说话人分离 | 5.0/10

 · 更新于 2026-09-10 · 约 17 分钟 · 8189 字 阅读 →