论文解读

SLAP: Scalable Language-Audio Pretraining with Variable-Duration Audio and Multi-Objective Training

音频检索 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4546 字 阅读 →
论文解读

SmoothCLAP: Soft-Target Enhanced Contrastive Language-Audio Pretraining for Affective Computing

语音情感识别 | 6.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3757 字 阅读 →
论文解读

SONAR: Self-Distilled Continual Pre-Training for Domain Adaptive Audio Representation

音频事件检测 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3901 字 阅读 →
论文解读

SPAM: Style Prompt Adherence Metric for Prompt-Based TTS

语音合成 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4628 字 阅读 →
论文解读

Speaking Clearly: A Simplified Whisper-Based Codec for Low-Bitrate Speech Coding

语音编码 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4729 字 阅读 →
论文解读

Speech Emotion Recognition based on Hierarchical Transformer with Shifted Windows

语音情感识别 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4324 字 阅读 →
论文解读

SpeechMapper: Speech-To-Text Embedding Projector for LLMs

语音大模型 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5397 字 阅读 →
论文解读

Syncspeech: Efficient and Low-Latency Text-to-Speech Based on Temporal Masked Transformer

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4853 字 阅读 →
论文解读

TAGARELA - A Portuguese Speech Dataset from Podcasts

语音识别 语音合成 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3967 字 阅读 →
论文解读

TASU: Text-only Alignment for Speech Understanding

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4344 字 阅读 →
论文解读

Test Time Adaptation for Speech Emotion Recognition

语音情感识别 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4724 字 阅读 →
论文解读

Text2Move: Text-To-Moving Sound Generation via Trajectory Prediction and Temporal Alignment

空间音频 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4225 字 阅读 →
论文解读

The 3rd Clarity Prediction Challenge: A Machine Learning Challenge for Hearing aid Speech Intelligibility Prediction

语音增强 | 7.5/10

 · 更新于 2026-09-25 · 约 7 分钟 · 3264 字 阅读 →
论文解读

The Synergistic Role of Audio and Large Video-Language Model in Source-Free Video Domain Adaptation

领域适应 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4434 字 阅读 →
论文解读

Thinking While Listening: Simple Test Time Scaling for Audio Classification

音频分类 | 6.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4620 字 阅读 →
论文解读

Timbre-Based Pretraining with Pseudo-Labels for Multi-Instrument Automatic Music Transcription

音乐信息检索 | 7.0/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7980 字 阅读 →
论文解读

Tldiffgan: A Latent Diffusion-Gan Framework with Temporal Information Fusion for Anomalous Sound Detection

音频事件检测 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4622 字 阅读 →
论文解读

Tpeformer: Temporal Patch Embedding Transformer

语音情感识别 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4746 字 阅读 →
论文解读

Training-Free Inference-Time Scaling for Audio Source Separation

语音增强 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3556 字 阅读 →
论文解读

Tri-Attention Fusion: Joint Temporal-Spectral and Bidirectional Modeling for Speech Spoofing Detection

语音伪造检测 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5146 字 阅读 →