论文解读

FlowFake: Liquid Networks for Audio Deepfake Detection

模型压缩 | 8.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6691 字 阅读 →
论文解读

Hybrid Diffusion Transformer for Instruction-Guided Audio Editing via Rectified Flow

Transformer | 7.6/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5452 字 阅读 →
论文解读

CraBERT: Efficient Phoneme Encoder Pre-Training via Cascade Fusion of Subword Representations for Text-to-Speech

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5796 字 阅读 →
论文解读

Probing Low Frame Rate Degradation in Neural Audio Codecs

语音生成 | 8.6/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5018 字 阅读 →
论文解读

Efficiency-Performance Trade-offs in Neural Speaker Diarization via Structured Pruning and Low-Bit Quantization

说话人日志 | 5.1/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6022 字 阅读 →
论文解读

Positional Encoding in the Context of Memristor-Based Analog Computation for Automatic Speech Recognition

语音识别 | 8/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4178 字 阅读 →
论文解读

Self-Guidance: Enhancing Neural Codecs via Decoder Manifold Alignment

语音合成 | 9.7/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6236 字 阅读 →
论文解读

Massive Open-Vocabulary Keyword Spotting

语音识别 | 9.8/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5180 字 阅读 →
论文解读

SpAArSIST: Sparsified AASIST for Efficient and Reliable Anti-Spoofing

模型压缩 | 7.7/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5272 字 阅读 →
论文解读

Towards Data-free and Training-free Compression for Speech Foundation Models Using Parameter Clustering

语音识别 | 6.4/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5326 字 阅读 →
论文解读

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling

语音识别 | 8.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5701 字 阅读 →
论文解读

On Low-Bit Quantization Errors in Speaker Verification: Diagnostic and Mitigation

说话人验证 | 6.6/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5313 字 阅读 →
论文解读

TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech

语音合成 | 8.2/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6310 字 阅读 →
论文解读

dots.tts Technical Report

语音合成 | 9/10

 · 更新于 2026-09-25 · 约 7 分钟 · 3032 字 阅读 →
论文解读

Spiking and Event-driven Neuromorphic Mamba Models for Efficient Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6692 字 阅读 →
论文解读

Data-Efficient On-Policy Distillation for Automatic Speech Recognition

语音识别 | 5.1/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4830 字 阅读 →
论文解读

HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding

语音合成 | 8.6/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6179 字 阅读 →
论文解读

TARQ: Tail-Aware Reconstruction Quantization for Rare-Word Robust Automatic Speech Recognition

语音识别 | 10/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6908 字 阅读 →
论文解读

Toward Native Multimodal Modeling: A Roadmap

多模态模型 | 10/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4994 字 阅读 →
论文解读

WaveNeXt 2: ConvNeXt-Based Fast Neural Vocoders With Residual Denoising and Sub-Modeling for GAN and Diffusion Models

语音合成 | 9.4/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5426 字 阅读 →