论文解读

SONAR: Self-Distilled Continual Pre-Training for Domain Adaptive Audio Representation

音频事件检测 | 7.0/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3901 字 阅读 →
论文解读

Sparse Autoencoders Make Audio Foundation Models More Explainable

模型评估 | 6.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4339 字 阅读 →
论文解读

Sparse-View Visual-Acoustic Latent Learning for Novel-View Audio Synthesis

空间音频 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5274 字 阅读 →
论文解读

Spatially Aware Self-Supervised Models for Multi-Channel Neural Speaker Diarization

说话人分离 | 8.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5315 字 阅读 →
论文解读

Speech Quality-Based Localization of Low-Quality Speech and Text-to-Speech Synthesis Artefacts

语音质量评估 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4591 字 阅读 →
论文解读

STACodec: Semantic Token Assignment for Balancing Acoustic Fidelity and Semantic Information in Audio Codecs

语音识别 | 8.0/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5581 字 阅读 →
论文解读

Temporal Distillation for Music Representation Learning

音乐信息检索 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4668 字 阅读 →
论文解读

Temporal Graph Modeling for Speech Emotion Recognition Using LSTM-Aggregated Multigraph Networks

语音情感识别 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3621 字 阅读 →
论文解读

Temporally Heterogeneous Graph Contrastive Learning for Multimodal Acoustic Event Classification

音频事件检测 | 8.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3902 字 阅读 →
论文解读

The Curious Case of Visual Grounding: Different Effects for Speech-and Text-Based Language Encoders

模型评估 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4508 字 阅读 →
论文解读

The Role of Prosodic and Lexical Cues in Turn-Taking with Self-Supervised Speech Representations

语音对话系统 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4558 字 阅读 →
论文解读

Timbre-Based Pretraining with Pseudo-Labels for Multi-Instrument Automatic Music Transcription

音乐信息检索 | 7.0/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7980 字 阅读 →
论文解读

TinyMU: A Compact Audio-Language Model for Music Understanding

音乐理解 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5011 字 阅读 →
论文解读

Toward Faithful Explanations in Acoustic Anomaly Detection

音频事件检测 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3867 字 阅读 →
论文解读

Towards Blind Data Cleaning: A Case Study in Music Source Separation

音乐信息检索 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4929 字 阅读 →
论文解读

Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6217 字 阅读 →
论文解读

Towards Lightweight Adaptation of Speech Enhancement Models in Real-World Environments

语音增强 | 8.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4520 字 阅读 →
论文解读

Understanding the Strengths and Weaknesses of SSL Models for Audio Deepfake Model Attribution

音频深度伪造检测 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4679 字 阅读 →
论文解读

Unsupervised Lexicon Learning from Speech is Limited by Representations Rather than Clustering

语音发现 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4486 字 阅读 →
论文解读

VBx for End-to-End Neural and Clustering-Based Diarization

说话人分离 | 8.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4816 字 阅读 →