论文解读

Staged Diffusion with Hybrid Mixture-of-Experts (MOE) for Multimodal Sentiment Analysis

语音情感识别 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4137 字 阅读 →
论文解读

Stemphonic: All-At-Once Flexible Multi-Stem Music Generation

音乐生成 | 7.7/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4942 字 阅读 →
论文解读

StereoFoley: Object-Aware Stereo Audio Generation from Video

音频生成 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3564 字 阅读 →
论文解读

Stereophonic Acoustic Echo Cancellation Using an Improved Affine Projection Algorithm with Adaptive Multiple Sub-Filters

语音增强 | 6.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4348 字 阅读 →
论文解读

Still Thinking or Stopped Talking? Dialogue Silence Intention Classification Using Multimodal Large Language Model

语音对话系统 | 6.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3789 字 阅读 →
论文解读

Str-DiffSep: Streamable Diffusion Model for Speech Separation

语音分离 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4807 字 阅读 →
论文解读

Stream-Voice-Anon: Enhancing Utility of Real-Time Speaker Anonymization Via Neural Audio Codec and Language Models

语音匿名化 | 7.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5467 字 阅读 →
论文解读

Streaming Speech Recognition with Decoder-Only Large Language Models and Latency Optimization

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4706 字 阅读 →
论文解读

Streamingbench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

基准测试 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3905 字 阅读 →
论文解读

StreamMark: A Deep Learning-Based Semi-Fragile Audio Watermarking for Proactive Deepfake Detection

音频深度伪造检测 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4723 字 阅读 →
论文解读

Stress Prediction from Temporal Emotion Trajectories in Clinical Patient-Physician Conversations

语音情感识别 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4349 字 阅读 →
论文解读

Structure-Aware Diffusion Schrödinger Bridge

数据集对齐 | 7.7/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4076 字 阅读 →
论文解读

StyHarmo: Efficient Style-Specific Video Generation with Music Synchronization

视频生成 | 6.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5066 字 阅读 →
论文解读

Style Attack Disguise: When Fonts Become a Camouflage for Adversarial Intent

对抗样本 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4288 字 阅读 →
论文解读

Style-Disentangled Diffusion for Controllable and Identity-Generalized Speech-Driven Body Motion Generation

语音驱动动作生成 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4058 字 阅读 →
论文解读

StyleBench: Evaluating Speech Language Models on Conversational Speaking Style Control

基准测试 | 8.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4238 字 阅读 →
论文解读

StylePitcher: Generating Style-Following and Expressive Pitch Curves for Versatile Singing Tasks

歌唱语音合成 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5251 字 阅读 →
论文解读

Subgraph Localization in the Subbands for Partially Spoofed Speech Detection

音频深度伪造检测 | 8.0/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3790 字 阅读 →
论文解读

Subsequence SDTW: Differentiable Alignment with Flexible Boundary Conditions

音乐信息检索 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4127 字 阅读 →
论文解读

Subspace Hybrid Adaptive Filtering for Phonocardiogram Signal Denoising

音频增强 | 7.0/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3758 字 阅读 →