论文解读

Sunac: Source-Aware Unified Neural Audio Codec

音频生成 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4495 字 阅读 →
论文解读

SURE: Synergistic Uncertainty-Aware Reasoning for Multimodal Emotion Recognition in Conversations

语音情感识别 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4997 字 阅读 →
论文解读

SwitchCodec: Adaptive Residual-Expert Sparse Quantization for High-Fidelity Neural Audio Coding

音频生成 | 8.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5387 字 阅读 →
论文解读

Symphony Rendering: Midi and Composer-Conditioned Auto Orchestration with Flow-Matching Transformers

音乐生成 | 7.0/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5626 字 阅读 →
论文解读

SynaSpot: A Lightweight, Streaming Multi-modal Framework for Keyword Spotting with Audio-Text Synergy

关键词检测 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3921 字 阅读 →
论文解读

Synchronous Secondary Path Modeling and Kronecker-Factorized Adaptive Algorithm for Multichannel Active Noise Control

主动噪声控制 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4919 字 阅读 →
论文解读

Syncspeech: Efficient and Low-Latency Text-to-Speech Based on Temporal Masked Transformer

语音合成 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4853 字 阅读 →
论文解读

SynParaSpeech: Automated Synthesis of Paralinguistic Datasets for Speech Generation and Understanding

语音合成 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4862 字 阅读 →
论文解读

Synthcloner: Synthesizer-Style Audio Transfer via Factorized Codec with ADSR Envelope Control

音频生成 | 8.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5185 字 阅读 →
论文解读

Synthesized Data Selection via Score Distribution Matching for Te Reo Māori Automatic Speech Recognition

语音识别 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4219 字 阅读 →
论文解读

Synthetic Data Domain Adaptation for ASR via LLM-Based Text and Phonetic Respelling Augmentation

语音识别 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4965 字 阅读 →
论文解读

Synthetic yet Striking? Assessing Vocal Charisma in TTS via Perceptual and Algorithmic Measures

语音合成 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4258 字 阅读 →
论文解读

T-Cache: Fast Inference For Masked Generative Transformer-Based TTS Via Prompt-Aware Feature Caching

语音合成 | 9.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4509 字 阅读 →
论文解读

T-Mimi: A Transformer-Based Mimi Decoder for Real-Time On-Phone TTS

语音合成 | 7.0/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3841 字 阅读 →
论文解读

TAG: Structured Temporal Audio Generation via LLM-Guided Manual Scription and Control

音频生成 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4086 字 阅读 →
论文解读

TAGARELA - A Portuguese Speech Dataset from Podcasts

语音识别 语音合成 | 7.0/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3967 字 阅读 →
论文解读

Taming Audio VAEs via Target-KL Regularization

音频生成 | 6.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5919 字 阅读 →
论文解读

Target Speaker Anonymization in Multi-Speaker Recordings

语音匿名化 | 7.6/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3887 字 阅读 →
论文解读

Target-Speaker LLM-ASR with Speaker-Aware Speech Encoder

语音识别 | 8.8/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4444 字 阅读 →
论文解读

Task Vector in TTS: Toward Emotionally Expressive Dialectal Speech Synthesis

语音合成 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4678 字 阅读 →