论文解读

Fast Speech Foundation Model Distillation Using Interleaved Stacking

知识蒸馏 | 6.6/10

 · 更新于 2026-09-08 · 约 10 分钟 · 4840 字 阅读 →
论文解读

Fast-SDE: Efficient Single-Microphone Sound Source Distance Estimation in Reverberant Environments

Fast-SDE: Efficient Single-Microphone Sound Source Distance Estimation in Reverberant Environments

 · 更新于 2026-09-08 · 约 12 分钟 · 5649 字 阅读 →
论文解读

Feature-Aligned Speech Watermarking for Robustness to Reconstruction Distortions

鲁棒性 | 7.1/10

 · 更新于 2026-09-08 · 约 15 分钟 · 7369 字 阅读 →
论文解读

Frozen Multimodal Embeddings for Personality and Cognitive Ability Assessment in Asynchronous Video Interviews

语音情感识别 | 6.7/10

 · 更新于 2026-09-08 · 约 13 分钟 · 6098 字 阅读 →
论文解读

Gumbel-BEARD: Automatic Layer Selection for Self-Supervised Adaptation of Whisper in Low-Resource Domains

语音识别 | 9.1/10

 · 更新于 2026-09-08 · 约 10 分钟 · 4665 字 阅读 →
论文解读

HALO: Half-Frame-Rate Adaptive Learnable Operator for Lightweight STFT-Based Speech Enhancement

语音增强 | 8.4/10

 · 更新于 2026-09-08 · 约 10 分钟 · 4830 字 阅读 →
论文解读

I Understand How You Feel: Enhancing Deeper Emotional Support Through Multilingual Emotional Validation in Dialogue System

语音识别 | 5.8/10

 · 更新于 2026-09-08 · 约 15 分钟 · 7227 字 阅读 →
论文解读

Interpreting and Steering a Text-to-Speech Language Model with Sparse Autoencoders

语音合成 | 7.7/10

 · 更新于 2026-09-08 · 约 11 分钟 · 5248 字 阅读 →
论文解读

Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip Synchronization

语音合成 | 6.8/10

 · 更新于 2026-09-08 · 约 12 分钟 · 5771 字 阅读 →
论文解读

Lung-SRAD: Spectral-Aware Regularized Audio DASS with Dual-Axis Patch-Mix Contrastive Learning for Respiratory Sound Classification

对比学习 | 6.8/10

 · 更新于 2026-09-08 · 约 11 分钟 · 5479 字 阅读 →
论文解读

MA-DLE: Speech-based Automatic Depression Level Estimation via Memory Augmentation

语音情感识别 | 7.5/10

 · 更新于 2026-09-08 · 约 12 分钟 · 5796 字 阅读 →
论文解读

Massive Open-Vocabulary Keyword Spotting

语音识别 | 9.8/10

 · 更新于 2026-09-08 · 约 11 分钟 · 5180 字 阅读 →
论文解读

Overcoming State Inertia in Full-Duplex Spoken Language Models via Activation Steering

基准测试 | 5.5/10

 · 更新于 2026-09-08 · 约 10 分钟 · 4609 字 阅读 →
论文解读

PianoKontext: Expressive Performance Rendering from Deadpan Context

音乐生成 | 9.1/10

 · 更新于 2026-09-08 · 约 9 分钟 · 4271 字 阅读 →
论文解读

Pretrained self-supervised speech models can recognize unseen consonants

语音识别 | 6.5/10

 · 更新于 2026-09-08 · 约 10 分钟 · 4929 字 阅读 →
论文解读

Quality Adaptive Angular Margin Learning for Respiratory Sound Classification

音频质量评估 | 9.5/10

 · 更新于 2026-09-08 · 约 18 分钟 · 8931 字 阅读 →
论文解读

RAIL: Rethinking Auditory Intelligence in Large Audio-Language Models with a CHC-Grounded Benchmark

音频问答 | 9.6/10

 · 更新于 2026-09-08 · 约 14 分钟 · 6896 字 阅读 →
论文解读

Real-Time Language Model Jamming: A Case Study for Live Music Accompaniment Generation

音乐信息检索 | 8.7/10

 · 更新于 2026-09-08 · 约 12 分钟 · 5795 字 阅读 →
论文解读

SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations

语音合成 | 7.9/10

 · 更新于 2026-09-08 · 约 10 分钟 · 4632 字 阅读 →
论文解读

Sensitivity Analysis of Generative Spatial Audio Metrics: A Study on Responsiveness, Smoothness, and Symmetry

音频生成 | 7.2/10

 · 更新于 2026-09-08 · 约 12 分钟 · 5919 字 阅读 →