论文解读

HARP: Harmonic-Aware Residual Partitioning for Neural Audio Codecs

音频编码 | 9.6/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6824 字 阅读 →
论文解读

How Reliable Are Multimodal Signals of Conversational State? Evidence from Remote Dyadic Collaborative Tasks

鲁棒性 | 6.6/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7971 字 阅读 →
论文解读

Is One Score Enough? Assessing Singing Quality of Songs with Temporal Score Curves

音乐理解 | 7.6/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7120 字 阅读 →
论文解读

Modeling turn-taking with distant viewing: investigating silence thresholds in human and AI-generated discourse

音视频 | 6.2/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6168 字 阅读 →
论文解读

Multi-Level Privacy-Preserving Dementia Detection from Speech via Targeted Adversarial Obfuscation and Representation Learning

语音属性识别 | 5.5/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6412 字 阅读 →
论文解读

NABEATs: Noise-Aware Audio Representation Learning

音频理解 | 6.7/10

 · 更新于 2026-09-06 · 约 18 分钟 · 8604 字 阅读 →
论文解读

Pseudo-label distillation for discriminative anomalous sound detection

音频事件检测 | 9.0/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7353 字 阅读 →
论文解读

Re-Sonance: A Dysarthric Asynchronous Real-Time Speech Conversion System Based on a Three-Stage Cascaded ASR-LLM-TTS Architecture

语音转换 | 6.9/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7105 字 阅读 →
论文解读

RealDESED: A Real-World Domestic Sound Event Detection Benchmark

音频事件检测 | 7.9/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7315 字 阅读 →
论文解读

Robust Summarization of Doctor-Patient Conversations: TalTech Systems for the Beyond Transcription Challenge

语音交互 | 6.3/10

 · 更新于 2026-09-06 · 约 18 分钟 · 8701 字 阅读 →
论文解读

SALMONN-2: Advancing General-Purpose Hearing Abilities with Self-Supervised Representations

音频理解 | 9.4/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7986 字 阅读 →
论文解读

Should Missing Modalities Always Be Necessary to Repair for Multi-modal Sentiment Analysis?

多模态模型 | 7.0/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7627 字 阅读 →
论文解读

SSTMark: Robust Training-Free Semantic-Level Speech Watermarking

音频水印 | 6.5/10

 · 更新于 2026-09-06 · 约 21 分钟 · 10235 字 阅读 →
论文解读

Team RAS in 11th ABAW Competition: Multimodal Ambivalence Recognition Approach

语音情感识别 | 5.3/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7032 字 阅读 →
论文解读

The tttAI System for the TSA-ASR Task of the SmartGlasses Challenge 2026

说话人日志 | 6.5/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6496 字 阅读 →
论文解读

Time-Frequency Consistency Learning for Robust Speech Deepfake Detection

语音伪造检测 | 7.9/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6790 字 阅读 →
论文解读

When to Use Extra Context: Evidence-Grounded Terminology Adaptation for Simultaneous Speech Translation

语音翻译 | 6.7/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6074 字 阅读 →
论文解读

X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System

语音翻译 | 7.8/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7793 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-07-21

共分析 34 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 118 分钟 · 58810 字 阅读 →
论文解读

A Geometry-Limited Identification Floor and Its Consequences for Voice-Clone Attribution in Professional Voice Actors

说话人验证 | 8.8/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8183 字 阅读 →