论文解读

NeuroSonic: Conditional Flow Matching for EEG-to-Speech Reconstruction

语音生成 | 7/10

 · 更新于 2026-09-07 · 约 13 分钟 · 6037 字 阅读 →
论文解读

ParaPairAudioBench: Paralinguistic Pairwise Audio Benchmark for LALM-as-a-Judge

语音质量评估 | 8.2/10

 · 更新于 2026-09-07 · 约 10 分钟 · 4943 字 阅读 →
论文解读

Perceptual Evaluation of Higher-Order Ambisonic Codecs on Both Synthetic Mixing and Native Recordings

音频编码 | 8/10

 · 更新于 2026-09-07 · 约 13 分钟 · 6403 字 阅读 →
论文解读

Poster: Exploring the Limits of Audio-Based Detection of Turkish Phone Call Scams

Poster: Exploring the Limits of Audio-Based Detection of Turkish Phone Call Scams

 · 更新于 2026-09-07 · 约 19 分钟 · 9355 字 阅读 →
论文解读

Progressive Alignment Objectives for Aligner-Encoder based ASR

语音识别 | 7.5/10

 · 更新于 2026-09-07 · 约 6 分钟 · 2719 字 阅读 →
论文解读

Real-Time Interactive Music Generation via Data-Free Streaming Consistency Distillation

音乐生成 | 7.1/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5781 字 阅读 →
论文解读

Selective Capability Unlearning in End-to-End Spoken Language Understanding

Selective Capability Unlearning in End-to-End Spoken Language Understanding

 · 更新于 2026-09-07 · 约 4 分钟 · 1826 字 阅读 →
论文解读

Sonus Health: Calibrated Heart-Murmur Detection from Smartphone-Based Veterinary Auscultation

音频事件检测 | 5.7/10

 · 更新于 2026-09-07 · 约 10 分钟 · 4683 字 阅读 →
论文解读

SphereVBx: Spherical Variational Bayes Clustering for Simplified EEND-VC Diarization

无监督学习 | 8.3/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5625 字 阅读 →
论文解读

Statistical validation and full-sphere extension of a Bayesian model for human static sound localisation

音频质量评估 | 6.7/10

 · 更新于 2026-09-07 · 约 16 分钟 · 7634 字 阅读 →
论文解读

Suppressing spectral edge effects in Schroeder Harmonic Complex

语音增强 | 7.3/10

 · 更新于 2026-09-07 · 约 10 分钟 · 4959 字 阅读 →
论文解读

The effect of micro-changes in the pluck trajectory on the sound of an acoustic guitar

信号处理基础 | 6.8/10

 · 更新于 2026-09-07 · 约 13 分钟 · 6466 字 阅读 →
论文解读

video-SALMONN-R\(^3\): Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding

多模态模型 | 8.2/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5751 字 阅读 →
论文解读

VieSpeaker: A Large-Scale Vietnamese Speaker Recognition Dataset Beyond Visual Dependency

说话人识别 | 7.5/10

 · 更新于 2026-09-07 · 约 10 分钟 · 4860 字 阅读 →
论文解读

ZONOS2 Technical Report

语音合成 | 10/10

 · 更新于 2026-09-07 · 约 15 分钟 · 7161 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-24

共分析 39 篇语音/AI 论文

 · 更新于 2026-09-07 · 约 98 分钟 · 48791 字 阅读 →
论文解读

A DDSP Framework for Adaptive Room Equalization

自适应滤波 | 6.5/10

 · 更新于 2026-09-07 · 约 13 分钟 · 6448 字 阅读 →
论文解读

A Generalized Formalism of Auto-Regressive Decoding for Speech Processing

自回归模型 | 4.1/10

 · 更新于 2026-09-07 · 约 10 分钟 · 4512 字 阅读 →
论文解读

Acoustic Landmark Detector based on Conformer and HuBERT

语音识别 | 5.5/10

 · 更新于 2026-09-07 · 约 21 分钟 · 10238 字 阅读 →
论文解读

Adding Robust Code-Switching Capabilities to High Performance Multilingual ASR

语音识别 | 7/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5481 字 阅读 →