论文解读

Speech Quality-Based Localization of Low-Quality Speech and Text-to-Speech Synthesis Artefacts

语音质量评估 | 7.0/10

 · 更新于 2026-09-16 · 约 10 分钟 · 4591 字 阅读 →
论文解读

SpeechCT-CLIP: Distilling Text-Image Knowledge to Speech for Voice-Native Multimodal CT Analysis

医疗AI | 7.5/10

 · 更新于 2026-09-16 · 约 9 分钟 · 4415 字 阅读 →
论文解读

SpeechMapper: Speech-To-Text Embedding Projector for LLMs

语音大模型 | 7.0/10

 · 更新于 2026-09-16 · 约 11 分钟 · 5397 字 阅读 →
论文解读

Spike-Driven Low-Power Speech Bandwidth Extension

语音增强 | 8.0/10

 · 更新于 2026-09-16 · 约 9 分钟 · 4404 字 阅读 →
论文解读

Spiking Attention Network: A Hybrid Neuromorphic Approach to Underwater Acoustic Localization and Zero-Shot Adaptation

声源定位 | 7.0/10

 · 更新于 2026-09-16 · 约 8 分钟 · 3981 字 阅读 →
论文解读

Spiking Temporal-Enhanced Network for Zero-Shot Audio-Visual Learning

音频分类 | 7.0/10

 · 更新于 2026-09-16 · 约 9 分钟 · 4341 字 阅读 →
论文解读

Spring Reverb Emulation with Hybrid Gated Convolutional Networks and State Space Models

音频生成 | 7.5/10

 · 更新于 2026-09-16 · 约 11 分钟 · 5227 字 阅读 →
论文解读

SSVD-O: Parameter-Efficient Fine-Tuning with Structured SVD for Speech Recognition

语音识别 | 7.0/10

 · 更新于 2026-09-16 · 约 11 分钟 · 5250 字 阅读 →
论文解读

ST-HNTM: Joint Speech-Text Neural Topic Modeling on the Hypersphere

主题建模 | 7.0/10

 · 更新于 2026-09-16 · 约 10 分钟 · 4767 字 阅读 →
论文解读

STACodec: Semantic Token Assignment for Balancing Acoustic Fidelity and Semantic Information in Audio Codecs

语音识别 | 8.0/10

 · 更新于 2026-09-16 · 约 12 分钟 · 5581 字 阅读 →
论文解读

Staged Diffusion with Hybrid Mixture-of-Experts (MOE) for Multimodal Sentiment Analysis

语音情感识别 | 8.0/10

 · 更新于 2026-09-16 · 约 9 分钟 · 4137 字 阅读 →
论文解读

Stemphonic: All-At-Once Flexible Multi-Stem Music Generation

音乐生成 | 7.7/10

 · 更新于 2026-09-16 · 约 10 分钟 · 4942 字 阅读 →
论文解读

Step-Audio-R1.5 Technical Report

语音对话系统 | 8.0/10

 · 更新于 2026-09-16 · 约 10 分钟 · 4640 字 阅读 →
论文解读

StereoFoley: Object-Aware Stereo Audio Generation from Video

音频生成 | 7.5/10

 · 更新于 2026-09-16 · 约 8 分钟 · 3564 字 阅读 →
论文解读

Stereophonic Acoustic Echo Cancellation Using an Improved Affine Projection Algorithm with Adaptive Multiple Sub-Filters

语音增强 | 6.0/10

 · 更新于 2026-09-16 · 约 9 分钟 · 4348 字 阅读 →
论文解读

Still Thinking or Stopped Talking? Dialogue Silence Intention Classification Using Multimodal Large Language Model

语音对话系统 | 6.5/10

 · 更新于 2026-09-16 · 约 8 分钟 · 3789 字 阅读 →
论文解读

Str-DiffSep: Streamable Diffusion Model for Speech Separation

语音分离 | 7.5/10

 · 更新于 2026-09-16 · 约 10 分钟 · 4807 字 阅读 →
论文解读

Stream-Voice-Anon: Enhancing Utility of Real-Time Speaker Anonymization Via Neural Audio Codec and Language Models

语音匿名化 | 7.0/10

 · 更新于 2026-09-16 · 约 11 分钟 · 5467 字 阅读 →
论文解读

Streaming Speech Recognition with Decoder-Only Large Language Models and Latency Optimization

语音识别 | 7.0/10

 · 更新于 2026-09-16 · 约 10 分钟 · 4706 字 阅读 →
论文解读

Streamingbench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

基准测试 | 7.5/10

 · 更新于 2026-09-16 · 约 8 分钟 · 3905 字 阅读 →