论文解读

Cumsum-Composable Phase Transport for Low-Cost Streaming Keyword Spotting

语音唤醒 | 5.9/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6381 字 阅读 →
论文解读

SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision

语音翻译 | 7.3/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7236 字 阅读 →
论文解读

Fretiq: Browser-Native Electric Guitar String Classification via Engineered Spectral Features and Held-Out Free-Play Evaluation

音频分类 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5587 字 阅读 →
论文解读

From a Multilingual Streaming ASR Backbone to Kenyan-Language Systems: Data-Centric Adaptation of Nemotron 3.5 for Kikuyu, Dholuo, and Kalenjin

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4601 字 阅读 →
论文解读

When to Use Extra Context: Evidence-Grounded Terminology Adaptation for Simultaneous Speech Translation

语音翻译 | 6.7/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6074 字 阅读 →
论文解读

Video = World + Event Stream

音频理解 | 4.9/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6888 字 阅读 →
论文解读

Do LLMs Need Architectural Changes for Simultaneous Speech Translation? A Prefix-to-Prefix Data Driven Approach

语音翻译 | 5.7/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7514 字 阅读 →
论文解读

Low-Latency Neural Models for Real-Time Music Enhancement

音乐源分离 | 7.7/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6872 字 阅读 →
论文解读

Breaking the Quality--Intelligibility Trade-off in Streaming Target Speaker Extraction via Deep-Feature-Anchored Preference Optimization

语音分离 | 6.3/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7612 字 阅读 →
论文解读

CoFi-Lite: Pushing the Limits of Ultra-Lightweight Speech Enhancement

语音增强 | 7.3/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6153 字 阅读 →
论文解读

LightMem-Ego: Your AI Memory for Everyday Life

流式处理 | 5.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5925 字 阅读 →
论文解读

The SonicAGI System for the REAL-TSE Challenge

语音分离 | 6.8/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8024 字 阅读 →
论文解读

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs

语音交互 | 9.2/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8435 字 阅读 →
论文解读

Metronome: Bound the Cache, Keep the Beat for Real-Time Interaction Model Serving

语音交互 | 8.7/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6175 字 阅读 →
论文解读

SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models

语音交互 | 8.9/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7382 字 阅读 →
论文解读

Streaming Neural Speech Codecs through Time-Invariant Representations

语音编码 | 6.0/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7723 字 阅读 →
论文解读

Wan-Streamer v0.2: Higher Resolution, Same Latency

音视频交互 | 5.4/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5455 字 阅读 →
论文解读

PhoStream: Benchmarking Real-World Streaming for Omnimodal Assistants in Mobile Scenarios

音视频问答 | 7.3/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7522 字 阅读 →
论文解读

ProactiveLLM: Learning Active Interaction for Streaming Large Language Models

语音识别 | 7.2/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3736 字 阅读 →
论文解读

Query-Based Asymmetric Modeling with Decoupled Input–Output Rates for Speech Restoration

语音增强 | 7.1/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6737 字 阅读 →