论文解读

RHO-PERFECT: Correlation Ceiling for Subjective Evaluation Datasets

模型评估 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4357 字 阅读 →
论文解读

S2Voice: Style-Aware Autoregressive Modeling with Enhanced Conditioning for Singing Style Conversion

歌唱语音转换 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5290 字 阅读 →
论文解读

SAASDNet: An EEG-Based Streaming Auditory Attention Switch Decoding Network for Self-Initiated Attention Switching in Mixed Speech

脑机接口 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4824 字 阅读 →
论文解读

Scalable Evaluation for Audio Identification Via Synthetic Latent Fingerprint Generation

音频检索 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3896 字 阅读 →
论文解读

Sing What You Fit: A Perception-Based Dataset and Benchmark for Vocal-Song Suitability Analysis

音乐信息检索 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 5002 字 阅读 →
论文解读

SingMOS-Pro: An Comprehensive Benchmark For Singing Quality Assessment

歌唱语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3984 字 阅读 →
论文解读

SP-MCQA: Evaluating Intelligibility of TTS Beyond the Word Level

语音合成 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4202 字 阅读 →
论文解读

SpeechCT-CLIP: Distilling Text-Image Knowledge to Speech for Voice-Native Multimodal CT Analysis

医疗AI | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4415 字 阅读 →
论文解读

Spring Reverb Emulation with Hybrid Gated Convolutional Networks and State Space Models

音频生成 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5227 字 阅读 →
论文解读

Still Thinking or Stopped Talking? Dialogue Silence Intention Classification Using Multimodal Large Language Model

语音对话系统 | 6.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3789 字 阅读 →
论文解读

StreamMark: A Deep Learning-Based Semi-Fragile Audio Watermarking for Proactive Deepfake Detection

音频深度伪造检测 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4723 字 阅读 →
论文解读

Symphony Rendering: Midi and Composer-Conditioned Auto Orchestration with Flow-Matching Transformers

音乐生成 | 7.0/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5626 字 阅读 →
论文解读

SymphonyGen: 3D Hierarchical Orchestral Generation with Controllable Harmony Skeleton

音乐生成 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5548 字 阅读 →
论文解读

SynParaSpeech: Automated Synthesis of Paralinguistic Datasets for Speech Generation and Understanding

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4862 字 阅读 →
论文解读

TAGARELA - A Portuguese Speech Dataset from Podcasts

语音识别 语音合成 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3967 字 阅读 →
论文解读

TAU: A Benchmark for Cultural Sound Understanding Beyond Semantics

音频问答 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3624 字 阅读 →
论文解读

Text2Move: Text-To-Moving Sound Generation via Trajectory Prediction and Temporal Alignment

空间音频 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4225 字 阅读 →
论文解读

The 3rd Clarity Prediction Challenge: A Machine Learning Challenge for Hearing aid Speech Intelligibility Prediction

语音增强 | 7.5/10

 · 更新于 2026-09-25 · 约 7 分钟 · 3264 字 阅读 →
论文解读

The Singing Voice Conversion Challenge 2025: From Singer Identity Conversion to Singing Style Conversion

歌唱语音转换 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3758 字 阅读 →
论文解读

The Structured Output Benchmark: A Multi-Source Benchmark for Evaluating Structured Output Quality in Large Language Models

基准测试 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3637 字 阅读 →