论文解读

I Hear, Therefore I Trust: A Socio-Technical Investigation of Humans as Synthetic Speech Detectors

语音合成 | 6.5/10

 · 更新于 2026-09-30 · 约 14 分钟 · 6689 字 阅读 →
论文解读

LoSATok: Low-dimensional Semantic-Acoustic Tokenizer for Cross-Domain Audio Understanding and Generation

语音合成 | 7/10

 · 更新于 2026-09-30 · 约 14 分钟 · 6912 字 阅读 →
论文解读

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation

语音生成 | 9.9/10

 · 更新于 2026-09-30 · 约 13 分钟 · 6224 字 阅读 →
论文解读

OmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher Distillation

音频检索 | 9.2/10

 · 更新于 2026-09-30 · 约 12 分钟 · 5784 字 阅读 →
论文解读

Robust Quantum-MUSIC for DoA Estimation Using Rydberg Atomic Receiver Arrays

Robust Quantum-MUSIC for DoA Estimation Using Rydberg Atomic Receiver Arrays

 · 更新于 2026-09-30 · 约 14 分钟 · 6718 字 阅读 →
论文解读

SMILE-Next: Teaching Large Language Models to Detect, Classify, and Reason about Laughter

语音情感识别 | 8.7/10

 · 更新于 2026-09-30 · 约 11 分钟 · 5028 字 阅读 →
论文解读

TARQ: Tail-Aware Reconstruction Quantization for Rare-Word Robust Automatic Speech Recognition

语音识别 | 10/10

 · 更新于 2026-09-30 · 约 14 分钟 · 6908 字 阅读 →
论文解读

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts

语音合成 | 8/10

 · 更新于 2026-09-30 · 约 10 分钟 · 4782 字 阅读 →
论文解读

Utilizing Missed Detections in Directional Sensitivity-Based DOA Estimation

语音识别 | 7.1/10

 · 更新于 2026-09-30 · 约 12 分钟 · 5920 字 阅读 →
论文解读

VoiceGiraffe: A Benchmark for Extreme Long-Context Audio-Language Understanding

音频问答 | 7.0/10

 · 更新于 2026-09-30 · 约 11 分钟 · 5379 字 阅读 →
论文解读

When Helpful Context Leaks: Privacy Risks in Domain-Adapted ASR

语音识别 | 10/10

 · 更新于 2026-09-30 · 约 13 分钟 · 6313 字 阅读 →
论文解读

Why We Need Speech to Evaluate Speech Translation

语音翻译 | 8.3/10

 · 更新于 2026-09-30 · 约 13 分钟 · 6500 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-28

共分析 30 篇语音/AI 论文

 · 更新于 2026-09-30 · 约 84 分钟 · 42014 字 阅读 →
论文解读

A Multimodal Framework for Dementia Detection via Linguistic and Acoustic Representation Learning

多模态模型 | 7.7/10

 · 更新于 2026-09-30 · 约 14 分钟 · 6784 字 阅读 →
论文解读

An investigation of AI integration in sound designer workflows and experiences

An investigation of AI integration in sound designer workflows and experiences

 · 更新于 2026-09-30 · 约 10 分钟 · 4636 字 阅读 →
论文解读

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models

多模态模型 | 9.7/10

 · 更新于 2026-09-30 · 约 13 分钟 · 6285 字 阅读 →
论文解读

Beyond Binary: Speech Representations Across the Cognitive Score Hierarchy

自监督学习 | 8.1/10

 · 更新于 2026-09-30 · 约 13 分钟 · 6023 字 阅读 →
论文解读

Can We Hear from Events? Generating Speech from Event Camera

语音合成 | 7.8/10

 · 更新于 2026-09-30 · 约 14 分钟 · 6885 字 阅读 →
论文解读

CFMDCTCodec: A Low-Bitrate Neural Speech Codec with Noise-Prior-aware Conditional Flow Matching for MDCT-Spectral Enhancement

语音编码 | 8.4/10

 · 更新于 2026-09-30 · 约 12 分钟 · 5946 字 阅读 →
论文解读

Continual Speaker Identity Unlearning with Minimal Interference

语音合成 | 8.3/10

 · 更新于 2026-09-30 · 约 4 分钟 · 1998 字 阅读 →