论文解读

AudioTQ: A Data-Oblivious 6-Bit CPU Audio Codec via Randomized Hadamard Rotation and Lloyd-Max Quantization

音频编码 | 3.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6965 字 阅读 →
论文解读

RT-SEMamba: Real-Time Speech Enhancement Mamba via Progressive Knowledge Distillation

语音增强 | 6.6/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6600 字 阅读 →
论文解读

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs

模型剪枝 | 8.4/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7856 字 阅读 →
论文解读

omni-macos: On-Device Omni-Modal Search on Apple Silicon

音频检索 | 7.2/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6772 字 阅读 →
论文解读

VoxZip: Semantic-Anchored Temporal KV Cache Compression for Long-Context Audio Inference

音频理解 | 7.7/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6697 字 阅读 →
论文解读

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models

音视频理解 | 7.7/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8473 字 阅读 →
论文解读

AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition

语音情感识别 | 5.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5579 字 阅读 →
论文解读

faster-enhancer.c: A Dependency-Free int8 Runtime for Streaming Speech Enhancement on Commodity CPUs

语音增强 | 8.7/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7387 字 阅读 →
论文解读

OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs

音视频问答 | 7.0/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9455 字 阅读 →
论文解读

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models

音视频问答 | 8.3/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8316 字 阅读 →
论文解读

VAD to the Bone: Ultra-Tiny Speech Activity Detection for Edge Deployment

语音活动检测 | 7.4/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6296 字 阅读 →
论文解读

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs

语音合成 | 7.6/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6369 字 阅读 →
论文解读

VibeVoice-ASR-BitNet Technical Report

语音识别 | 7.8/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6193 字 阅读 →
论文解读

Scalable Keyword Spotting via Modular Network Expansion

语音唤醒 | 7.1/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7687 字 阅读 →
论文解读

Ultra-Compact CNN Architectures for Tropical Bird Audio Detection on Microcontrollers

音频事件检测 | 9.3/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7444 字 阅读 →
论文解读

Staged Depth-Pruning Distillation of a Flow-Matching Text-to-Speech Teacher: A Compact Hindi Speech Synthesizer

语音合成 | 7.9/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8461 字 阅读 →
论文解读

Efficient Audio-Visual Event Recognition via Knowledge Distillation and Dynamic INT8 Quantization of a Hybrid Cross-Attention Network

音视频理解 | 5.8/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5753 字 阅读 →
论文解读

Do We Really Need Multimodal Emotion Language Models Larger Than 1B Parameters?

语音情感识别 | 7.4/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6817 字 阅读 →
论文解读

CoFi-Lite: Pushing the Limits of Ultra-Lightweight Speech Enhancement

语音增强 | 7.3/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6153 字 阅读 →
论文解读

LightMem-Ego: Your AI Memory for Everyday Life

流式处理 | 5.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5925 字 阅读 →