论文解读

Reasoning-Oriented Post-Training and Inference-Time LoRA Rescaling for Audio-Dependent Question Answering

音频理解 | 8.3/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5275 字 阅读 →
论文解读

sanoTTS: The Smallest Real-Time Neural TTS on a General-Purpose Microcontroller

语音合成 | 10.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4848 字 阅读 →
论文解读

Self-Supervised Speech Representations Track Spoken Language Convergence to Adult Models in Infants and Children Who Are Deaf/Hard-of-Hearing

语音属性识别 | 9.7/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4955 字 阅读 →
论文解读

Separating Voice from Age in COPD Screening

语音属性识别 | 7.2/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5333 字 阅读 →
论文解读

Simulation-to-Real First-Break Segmentation for Efficient Inversion in Musculoskeletal Ultrasound Tomography

音频事件检测 | 9.8/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4588 字 阅读 →
论文解读

SlimDiffuSE: Towards Efficient Diffusion-Based Speech Enhancement using Slimmable Networks

语音增强 | 7.6/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4594 字 阅读 →
论文解读

Spiking Neural Networks for Energy-Efficient Object Detection in Forward-Looking Sonar Imagery

音频事件检测 | 8.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4625 字 阅读 →
论文解读

TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming

音视频理解 | 10.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4537 字 阅读 →
论文解读

Towards Actionable Surgical Team Dynamics: from Teamwork to Counterfactual Annotations

音视频理解 | 8.9/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4656 字 阅读 →
论文解读

Training DeepFilterNet with Accurate Room Acoustic Simulations Improves Single-Channel Speech Enhancement

语音增强 | 9.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4932 字 阅读 →
论文解读

TurboBias 2.0: Streaming Context-Biasing for Production-Efficient ASR Systems

语音识别 | 9.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5444 字 阅读 →
论文解读

Unsupervised Speech Recognition at the Syllable Level

语音识别 | 8.9/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5121 字 阅读 →
论文解读

Vibrato Matching for Modulation Control and Blending in Sound Mixtures

音频生成 | 8.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4279 字 阅读 →
论文解读

WnW: Waxing-and-Waning KV Cache for Long-Form Speech LLMs

语音交互 | 9.4/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5062 字 阅读 →
论文解读

μNet: Ultra-Low-Memory and Low-Complexity Speech Enhancement for Embedded Digital Signal Processors

语音增强 | 8.3/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5258 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-08-25

共分析 46 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 101 分钟 · 50381 字 阅读 →
论文解读

A Resource-Efficient CNN-Based EEG Auditory Attention Decoding ASIC

助听器 | 6.2/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5899 字 阅读 →
论文解读

A Speech Corpus for Mizo Automatic Speech Recognition: Whisper and SraVaani 1.0 Fine-Tuning with Morphology-Aware Evaluation

语音识别 | 6.7/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5678 字 阅读 →
论文解读

Dancing Through Soundscapes: Designing a Low-Cost, Sound-Based Device for Sensing and Interpreting Movement and Dance

音频交互 | 5.2/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5265 字 阅读 →
论文解读

DAVSS: Distilled Audio-Visual State Space Models

音视频理解 | 6.4/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5722 字 阅读 →