论文解读

Towards Data-free and Training-free Compression for Speech Foundation Models Using Parameter Clustering

语音识别 | 6.4/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5326 字 阅读 →
论文解读

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling

语音识别 | 8.6/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5701 字 阅读 →
论文解读

On Low-Bit Quantization Errors in Speaker Verification: Diagnostic and Mitigation

说话人验证 | 6.6/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5313 字 阅读 →
论文解读

TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech

语音合成 | 8.2/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6310 字 阅读 →
论文解读

dots.tts Technical Report

语音合成 | 9/10

 · 更新于 2026-09-06 · 约 7 分钟 · 3032 字 阅读 →
论文解读

Spiking and Event-driven Neuromorphic Mamba Models for Efficient Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6692 字 阅读 →
论文解读

Data-Efficient On-Policy Distillation for Automatic Speech Recognition

语音识别 | 5.1/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4830 字 阅读 →
论文解读

HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding

语音合成 | 8.6/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6179 字 阅读 →
论文解读

TARQ: Tail-Aware Reconstruction Quantization for Rare-Word Robust Automatic Speech Recognition

语音识别 | 10/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6908 字 阅读 →
论文解读

Toward Native Multimodal Modeling: A Roadmap

多模态模型 | 10/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4994 字 阅读 →
论文解读

WaveNeXt 2: ConvNeXt-Based Fast Neural Vocoders With Residual Denoising and Sub-Modeling for GAN and Diffusion Models

语音合成 | 9.4/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5426 字 阅读 →
论文解读

Perforated Neural Networks for Keyword Spotting

关键词检测 | 5/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7432 字 阅读 →
论文解读

OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models

音视频 | 7.0/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8379 字 阅读 →
论文解读

Entropy-Monitored Kernelized Token Distillation for Audio-Visual Compression

音视频事件检测 | 8.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5068 字 阅读 →
论文解读

Cross-Architecture Knowledge Distillation of WavLM for Lightweight Speaker Verification

说话人验证 | 8.0/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6123 字 阅读 →
论文解读

Enhancing Speaker Verification with w2v-BERT 2.0 and Knowledge Distillation Guided Structured Pruning

说话人验证 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4903 字 阅读 →
论文解读

Lightweight and Generalizable Acoustic Scene Representations Via Contrastive Fine-Tuning and Distillation

音频场景理解 | 8.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5226 字 阅读 →
论文解读

MAGE: A Coarse-to-Fine Speech Enhancer with Masked Generative Model

语音增强 | 8.0/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5811 字 阅读 →
论文解读

S-SONDO: Self-Supervised Knowledge Distillation for General Audio Foundation Models

音频分类 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4416 字 阅读 →
论文解读

Transferable Audio Lottery Tickets: Gradient Accumulation for Extreme Sparsity

音频分类 | 7.0/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3873 字 阅读 →