论文解读

Task-Oriented Sound Privacy Preservation for Sound Event Detection Via End-to-End Adversarial Multi-Task Learning

音频事件检测 | 7.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5739 字 阅读 →
论文解读

TASU: Text-only Alignment for Speech Understanding

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4344 字 阅读 →
论文解读

TAU: A Benchmark for Cultural Sound Understanding Beyond Semantics

音频问答 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3624 字 阅读 →
论文解读

Teacher-Guided Pseudo Supervision and Cross-Modal Alignment for Audio-Visual Video Parsing

音视频 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4812 字 阅读 →
论文解读

Teaching Audio Models to Reason: A Unified Framework for Source- and Layer-Wise Distillation

音频问答 | 7.0/10

 · 更新于 2026-09-06 · 约 7 分钟 · 3417 字 阅读 →
论文解读

Teaching the Teachers: Boosting Unsupervised Domain Adaptation In Speech Recognition By Ensemble Update

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4428 字 阅读 →
论文解读

Temporal Distillation for Music Representation Learning

音乐信息检索 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4668 字 阅读 →
论文解读

Temporal Graph Modeling for Speech Emotion Recognition Using LSTM-Aggregated Multigraph Networks

语音情感识别 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3621 字 阅读 →
论文解读

Temporal-Spatial Decouple Before Act: Disentangled Representation Learning for Multimodal Sentiment Analysis

情感分析 | 7.5/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6191 字 阅读 →
论文解读

Temporally Heterogeneous Graph Contrastive Learning for Multimodal Acoustic Event Classification

音频事件检测 | 8.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3902 字 阅读 →
论文解读

Test Time Adaptation for Speech Emotion Recognition

语音情感识别 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4724 字 阅读 →
论文解读

Test-Time Scaling for Auditory Cognition in Audio Language Models

音频问答 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4888 字 阅读 →
论文解读

Testing The Efficient Coding Hypothesis Beyond Humans: The Auditory Kernels of Bat Vocalizations

生物声学 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3800 字 阅读 →
论文解读

Text2midi-InferAlign: Improving Symbolic Music Generation with Inference-Time Alignment

音乐生成 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3668 字 阅读 →
论文解读

Text2Move: Text-To-Moving Sound Generation via Trajectory Prediction and Temporal Alignment

空间音频 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4225 字 阅读 →
论文解读

TextlessRAG: End-to-End Visual Document RAG by Speech without Text

语音问答 | 8.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4218 字 阅读 →
论文解读

The 3rd Clarity Prediction Challenge: A Machine Learning Challenge for Hearing aid Speech Intelligibility Prediction

语音增强 | 7.5/10

 · 更新于 2026-09-06 · 约 7 分钟 · 3264 字 阅读 →
论文解读

The Curious Case of Visual Grounding: Different Effects for Speech-and Text-Based Language Encoders

模型评估 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4508 字 阅读 →
论文解读

The Impact of Audio Watermarking on Audio Anti-Spoofing Countermeasures

音频深度伪造检测 | 8.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4839 字 阅读 →
论文解读

The Muse Benchmark: Probing Music Perception and Auditory Relational Reasoning in Audio LLMs

音乐理解 | 8.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3655 字 阅读 →