论文解读

MVEB: Massive Video Embedding Benchmark

基准测试 | 6.5/10

 · 更新于 2026-09-07 · 约 15 分钟 · 7098 字 阅读 →
论文解读

Next-Turn: Duration-Aware Streaming Endpoint Detection via Time-to-Next-Speech-Onset Prediction

语音合成 | 7.9/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5858 字 阅读 →
论文解读

Non-Autoregressive Minimum Bayes' Risk Decoding for Fast Speech Recognition

Non-Autoregressive Minimum Bayes' Risk Decoding for Fast Speech Recognition

 · 更新于 2026-09-07 · 约 23 分钟 · 11197 字 阅读 →
论文解读

OlfactProfile: Profile-Conditioned Odor Prediction from Audiovisual Content

多模态模型 | 5.6/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5084 字 阅读 →
论文解读

One-Step Token-to-Waveform Generation with MeanFlow in Latent Space

语音合成 | 9.3/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5638 字 阅读 →
论文解读

Perceptual compensation for tonal context in self-supervised speech models

语音识别 | 7.7/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5350 字 阅读 →
论文解读

PhASE-Flow: Phonetic-Conditioned Acoustic Flow Matching in SSL Representation Domain for Speech Enhancement

语音增强 | 7.6/10

 · 更新于 2026-09-07 · 约 13 分钟 · 6160 字 阅读 →
论文解读

Reading between the Lines: Leveraging Large Language Models for Global Dementia and Depression Assessment from Clinical Interviews

语音情感识别 | 6.8/10

 · 更新于 2026-09-07 · 约 28 分钟 · 13976 字 阅读 →
论文解读

Single frequency filtering based multi-speaker direction of arrival estimation from stereo recordings

语音增强 | 7/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5711 字 阅读 →
论文解读

SpeechDx: A Multi-Task Benchmark for Clinical Speech AI

语音识别 | 7.6/10

 · 更新于 2026-09-07 · 约 14 分钟 · 6566 字 阅读 →
论文解读

Synergizing Zero-Shot Cross-Lingual Alzheimer Detection with Language-Invariant Multimodal Bi-Geometric Adversarial Learning

Synergizing Zero-Shot Cross-Lingual Alzheimer Detection with Language-Invariant Multimodal Bi-Geometric Adversarial Learning

 · 更新于 2026-09-07 · 约 13 分钟 · 6353 字 阅读 →
论文解读

Transductive Zero-Shot Audio Classification with Audio-Language Models

音频分类 | 6.4/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5272 字 阅读 →
论文解读

Turning music identification into a neural forward pass

音频分类 | 7.4/10

 · 更新于 2026-09-07 · 约 13 分钟 · 6103 字 阅读 →
论文解读

Vibrato Expression Control for Singing Voice Conversion with Improving Independent Control

Vibrato Expression Control for Singing Voice Conversion with Improving Independent Control

 · 更新于 2026-09-07 · 约 13 分钟 · 6110 字 阅读 →
论文解读

When Multiple Scripts Matter: Evaluating ASR in Clinical Settings

语音识别 | 9.1/10

 · 更新于 2026-09-07 · 约 13 分钟 · 6513 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-17

共分析 35 篇语音/AI 论文

 · 更新于 2026-09-07 · 约 94 分钟 · 46631 字 阅读 →
论文解读

Acoustic Prompting via Stage-wise Modulation for Few-Shot Learning in Audio Language Models

音频分类 | 8.3/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5504 字 阅读 →
论文解读

Acoustic, VOC, and Multimodal Stress Source Localization in the Internet of Plants

声源定位 | 9.7/10

 · 更新于 2026-09-07 · 约 15 分钟 · 7058 字 阅读 →
论文解读

AdaTT: Text-Guided Instrument Timbre Transfer with Target-Adaptive Structural Control

音频生成 | 8.7/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5380 字 阅读 →
论文解读

An Asymmetric Formula for Interval Consonance and its Relation to Harmonic Coincidence

音乐信息检索 | 8.0/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5496 字 阅读 →