论文解读

RCAL: Reinforced Cross-Modal Alignment for Multimodal Sentiment Analysis with Sparse Visual Frames

多模态模型 | 8.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4562 字 阅读 →
论文解读

Representation-Based Data Quality Audits for Audio

数据集 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4054 字 阅读 →
论文解读

Representation-Diverse Self-Supervision for Cross-Domain Bioacoustic Learning in Low-Resource Settings

生物声学 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5415 字 阅读 →
论文解读

Rethinking Entity Disambiguation in Complex Modalities

实体消歧 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4884 字 阅读 →
论文解读

Salad-VAE: Semantic Audio Compression with Language-Audio Distillation

音频压缩 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4645 字 阅读 →
论文解读

Semantic-Guided Pseudo-Feature Attention Network for Audio-Visual Zero-Shot Learning

音频分类 零样本学习 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4921 字 阅读 →
论文解读

SLAP: Scalable Language-Audio Pretraining with Variable-Duration Audio and Multi-Objective Training

音频检索 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4546 字 阅读 →
论文解读

SmoothCLAP: Soft-Target Enhanced Contrastive Language-Audio Pretraining for Affective Computing

语音情感识别 | 6.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3757 字 阅读 →
论文解读

SPAM: Style Prompt Adherence Metric for Prompt-Based TTS

语音合成 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4628 字 阅读 →
论文解读

Spatial-CLAP: Learning Spatially-Aware Audio–Text Embeddings for Multi-Source Conditions

空间音频 | 8.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4905 字 阅读 →
论文解读

Speech Emotion Recognition based on Hierarchical Transformer with Shifted Windows

语音情感识别 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4324 字 阅读 →
论文解读

SpeechCT-CLIP: Distilling Text-Image Knowledge to Speech for Voice-Native Multimodal CT Analysis

医疗AI | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4415 字 阅读 →
论文解读

Style-Disentangled Diffusion for Controllable and Identity-Generalized Speech-Driven Body Motion Generation

语音驱动动作生成 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4058 字 阅读 →
论文解读

SynaSpot: A Lightweight, Streaming Multi-modal Framework for Keyword Spotting with Audio-Text Synergy

关键词检测 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3921 字 阅读 →
论文解读

Temporally Heterogeneous Graph Contrastive Learning for Multimodal Acoustic Event Classification

音频事件检测 | 8.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3902 字 阅读 →
论文解读

The Curious Case of Visual Grounding: Different Effects for Speech-and Text-Based Language Encoders

模型评估 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4508 字 阅读 →
论文解读

Towards Effective Negation Modeling in Joint Audio-Text Models for Music

音乐理解 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4421 字 阅读 →
论文解读

TTA: Transcribe, Translate and Alignment for Cross-Lingual Speech Representation

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5261 字 阅读 →
论文解读

WavLink: Compact Audio–Text Embeddings with a Global Whisper Token

音频检索 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4681 字 阅读 →
论文解读

Robust Audio-Text Retrieval via Cross-Modal Attention and Hybrid Loss

音频检索 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4499 字 阅读 →