论文解读

Speech-based Psychological Crisis Assessment using LLMs

语音情感识别 | 5.8/10

 · 更新于 2026-09-10 · 约 17 分钟 · 8451 字 阅读 →
论文解读

Sub-JEPA: Subspace Gaussian Regularization for Stable End-to-End World Models

世界模型 | 5.0/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4662 字 阅读 →
论文解读

Towards Trustworthy Audio Deepfake Detection: A Systematic Framework for Diagnosing and Mitigating Gender Bias

音频深度伪造检测 | 6.5/10

 · 更新于 2026-09-10 · 约 18 分钟 · 8536 字 阅读 →
论文解读

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation

音视频生成 | 6.5/10

 · 更新于 2026-09-10 · 约 19 分钟 · 9059 字 阅读 →
论文解读

Voice Biomarkers for Depression and Anxiety

语音生物标志物 | 1.0/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4782 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-12

共分析 39 篇语音/AI 论文

 · 更新于 2026-09-10 · 约 157 分钟 · 78445 字 阅读 →
论文解读

A Decomposed Retrieval-Edit-Rerank Framework for Chord Generation

A Decomposed Retrieval-Edit-Rerank Framework for Chord Generation

 · 更新于 2026-09-10 · 约 15 分钟 · 7086 字 阅读 →
论文解读

Adaptive Regularization for Sparsity Control in Bregman-Based Optimizers

Adaptive Regularization for Sparsity Control in Bregman-Based Optimizers

 · 更新于 2026-09-10 · 约 16 分钟 · 7952 字 阅读 →
论文解读

Anisotropic Modality Align

Anisotropic Modality Align

 · 更新于 2026-09-10 · 约 16 分钟 · 7980 字 阅读 →
论文解读

Asymmetric Phase Coding Audio Watermarking

Asymmetric Phase Coding Audio Watermarking

 · 更新于 2026-09-10 · 约 18 分钟 · 8875 字 阅读 →
论文解读

BeeVe: Unsupervised Acoustic State Discovery in Honey Bee Buzzing

BeeVe: Unsupervised Acoustic State Discovery in Honey Bee Buzzing

 · 更新于 2026-09-10 · 约 20 分钟 · 9725 字 阅读 →
论文解读

Dependence on Early and Late Reverberation of Single-Channel Speaker Distance Estimation

Dependence on Early and Late Reverberation of Single-Channel Speaker Distance Estimation

 · 更新于 2026-09-10 · 约 13 分钟 · 6123 字 阅读 →
论文解读

Do Joint Audio-Video Generation Models Understand Physics?

Do Joint Audio-Video Generation Models Understand Physics?

 · 更新于 2026-09-10 · 约 18 分钟 · 8531 字 阅读 →
论文解读

Evaluating voice anonymisation using similarity rank disclosure

Evaluating voice anonymisation using similarity rank disclosure

 · 更新于 2026-09-10 · 约 15 分钟 · 7465 字 阅读 →
论文解读

MIST: Multimodal Interactive Speech-based Tool-calling Conversational Assistants for Smart Homes

MIST: Multimodal Interactive Speech-based Tool-calling Conversational Assistants for Smart Homes

 · 更新于 2026-09-10 · 约 14 分钟 · 7004 字 阅读 →
论文解读

Sparse Autoencoders as Plug-and-Play Firewalls for Adversarial Attack Detection in VLMs

Sparse Autoencoders as Plug-and-Play Firewalls for Adversarial Attack Detection in VLMs

 · 更新于 2026-09-10 · 约 18 分钟 · 8609 字 阅读 →
论文解读

TARNet: A Temporal-Aware Multi-Scale Architecture for Closed-Set Speaker Identification

TARNet: A Temporal-Aware Multi-Scale Architecture for Closed-Set Speaker Identification

 · 更新于 2026-09-10 · 约 18 分钟 · 8941 字 阅读 →
论文解读

Zero-Shot Imagined Speech Decoding via Imagined-to-Listened MEG Mapping

Zero-Shot Imagined Speech Decoding via Imagined-to-Listened MEG Mapping

 · 更新于 2026-09-10 · 约 13 分钟 · 6248 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-11

共分析 12 篇语音/AI 论文

 · 更新于 2026-09-10 · 约 41 分钟 · 20158 字 阅读 →
论文解读

Audio-Visual Intelligence in Large Foundation Models

跨模态 | 8.0/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4580 字 阅读 →