论文解读

Selective Hub Fusion with Modality-Heterogeneous Experts for Multimodal Emotion Recognition

多模态模型 | 6.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4869 字 阅读 →
论文解读

Self-Supervised Note Tracking and Multi-Pitch Estimation Via Reconstruction-Based Learning

多音高估计 音符跟踪 | 8.5/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9289 字 阅读 →
论文解读

Semantic Anchor Transfer from Short to Long Speech in a Distillation-Based Summarization Framework

语音摘要 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5537 字 阅读 →
论文解读

Semantic-Guided Pseudo-Feature Attention Network for Audio-Visual Zero-Shot Learning

音频分类 零样本学习 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4921 字 阅读 →
论文解读

SEP-ST: Incorporating Speech Entity Prompt Into Large Language Models for Speech Translation

语音翻译 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3975 字 阅读 →
论文解读

Separate this, and all of these Things Around It: Music Source Separation Via Hyperellipsoidal Queries

音乐分离 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4838 字 阅读 →
论文解读

Sequence-Level Unsupervised Training in Speech Recognition: A Theoretical Study

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 7 分钟 · 3356 字 阅读 →
论文解读

Sequential and Simultaneous Optimization of Microphone Array Geometry and Region-of-Interest Beamforming

声源定位 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4818 字 阅读 →
论文解读

Session-Level Spoken Language Assessment with A Multimodal Foundation Model Via Multi-Target Learning

语音评估 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4959 字 阅读 →
论文解读

SFM-TTS: Lightweight and Rapid Speech Synthesis with Flexible Shortcut Flow Matching

语音合成 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5280 字 阅读 →
论文解读

Shared Representation Learning for Reference-Guided Targeted Sound Detection

音频事件检测 | 8.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4764 字 阅读 →
论文解读

Shortcut Flow Matching for Speech Enhancement: Step-Invariant Flows via Single Stage Training

语音增强 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4631 字 阅读 →
论文解读

Sidon: Fast and Robust Open-Source Multilingual Speech Restoration for Large-Scale Dataset Cleansing

语音增强 | 8.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4546 字 阅读 →
论文解读

SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models

音频问答 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5286 字 阅读 →
论文解读

Sing What You Fit: A Perception-Based Dataset and Benchmark for Vocal-Song Suitability Analysis

音乐信息检索 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 5002 字 阅读 →
论文解读

Sing2Song: An Accompaniment Generation System Based on Solo Singing

音乐生成 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4905 字 阅读 →
论文解读

Single-Microphone Audio Point Source Discriminative Localization from Reverberation Late Tail Estimation

说话人分离 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3846 字 阅读 →
论文解读

Single-Step Controllable Music Bandwidth extension with Flow Matching

音乐信息检索 | 7.0/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5522 字 阅读 →
论文解读

SingMOS-Pro: An Comprehensive Benchmark For Singing Quality Assessment

歌唱语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3984 字 阅读 →
论文解读

SIREN: Spatially-Informed Reconstruction of Binaural Audio with Vision

空间音频 | 7.0/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6276 字 阅读 →