论文解读

FocalCodec-Stream: Streaming Low-Bitrate Speech Coding via Causal Distillation

语音编码 | 8.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5230 字 阅读 →
论文解读

IBPCodec : A Low-Bitrate Lightweight Speech Codec With Inter-Band Prediction

语音编码 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3768 字 阅读 →
论文解读

Int-MeanFlow: Few-Step Speech Generation with Integral Velocity Distillation

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5048 字 阅读 →
论文解读

Integrating Speaker Embeddings and LLM-Derived Semantic Representations for Streaming Speaker Diarization

说话人分离 | 6.5/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8207 字 阅读 →
论文解读

Lightweight Phoneme-Conditioned Bandwidth Extension for Body-Conducted Speech

语音增强 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4701 字 阅读 →
论文解读

Low-Bandwidth High-Fidelity Speech Transmission with Generative Latent Joint Source-Channel Coding

语音增强 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4889 字 阅读 →
论文解读

MeanVC: Lightweight and Streaming Zero-Shot Voice Conversion via Mean Flows

语音转换 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5574 字 阅读 →
论文解读

Online Register For Dual-Mode Self-Supervised Speech Models: Mitigating the Lack of Future Context

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5121 字 阅读 →
论文解读

Phrased: Phrase Dictionary Biasing for Speech Translation

语音翻译 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4074 字 阅读 →
论文解读

Real-Time Streaming MEL Vocoding with Generative Flow Matching

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4120 字 阅读 →
论文解读

SAASDNet: An EEG-Based Streaming Auditory Attention Switch Decoding Network for Self-Initiated Attention Switching in Mixed Speech

脑机接口 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4824 字 阅读 →
论文解读

SpatialNet-Echo: Real-Time Acoustic Echo Cancellation via Integrated Narrow-Band and Cross-Band Processing

语音增强 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4681 字 阅读 →
论文解读

Spike-Driven Low-Power Speech Bandwidth Extension

语音增强 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4404 字 阅读 →
论文解读

Str-DiffSep: Streamable Diffusion Model for Speech Separation

语音分离 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4807 字 阅读 →
论文解读

Streaming Speech Recognition with Decoder-Only Large Language Models and Latency Optimization

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4706 字 阅读 →
论文解读

SynaSpot: A Lightweight, Streaming Multi-modal Framework for Keyword Spotting with Audio-Text Synergy

关键词检测 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3921 字 阅读 →
论文解读

Syncspeech: Efficient and Low-Latency Text-to-Speech Based on Temporal Masked Transformer

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4853 字 阅读 →
论文解读

Train Short, Infer Long: Speech-LLM Enables Zero-Shot Streamable Joint ASR and Diarization on Long Audio

说话人分离 | 9.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4925 字 阅读 →
论文解读

VChangeCodec: An Ultra Low-Complexity Neural Speech Codec with Built-In Voice Changer for Customized Real-Time Communication

语音转换 语音增强 | 8.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5237 字 阅读 →
论文解读

VoXtream: Full-Stream Text-To-Speech With Extremely Low Latency

语音合成 | 8.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6227 字 阅读 →