论文解读

CoVA: Text-Guided Composed Video Retrieval for Audio-Visual Content

跨模态检索 | 6.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4170 字 阅读 →
论文解读

Cross-Modal Bottleneck Fusion for Noise Robust Audio-Visual Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4237 字 阅读 →
论文解读

DepthTalk: Few-Shot Talking Head Generation with Depth-Aware 3D Gaussian Field Motion

说话人生成 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3888 字 阅读 →
论文解读

Efficient Audio-Visual Inference Via Token Clustering And Modality Fusion

音频问答 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4866 字 阅读 →
论文解读

FastAV: Efficient Token Pruning for Audio-Visual Large Language Model Inference

音频问答 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4320 字 阅读 →
论文解读

FoleyBench: A Benchmark for Video-to-Audio Models

音频生成 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4402 字 阅读 →
论文解读

GMS-CAVP: Improving Audio-Video Correspondence with Multi-Scale Constrative and Generative Pretraining

音频生成 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3524 字 阅读 →
会议任务专题

ICASSP 2026 - 音视频

共 6 篇 ICASSP 2026 音视频 方向论文

 · 更新于 2026-09-25 · 约 18 分钟 · 8635 字 阅读 →
论文解读

Meanflow-Accelerated Multimodal Video-to-Audio Synthesis Via One-Step Generation

音频生成 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4951 字 阅读 →
论文解读

Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4158 字 阅读 →
论文解读

Motionbeat: Motion-Aligned Music Representation via Embodied Contrastive Learning and Bar-Equivariant Contact-Aware Encoding

舞蹈生成 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4092 字 阅读 →
论文解读

MSCT: Differential Cross-Modal Attention for Deepfake Detection

音频深度伪造检测 | 6.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3968 字 阅读 →
论文解读

Multimodal Self-Attention Network with Temporal Alignment for Audio-Visual Emotion Recognition

语音情感识别 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4714 字 阅读 →
论文解读

Noise-Robust AV-ASR Using Visual Features both in the Whisper Encoder and Decoder

语音识别 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4647 字 阅读 →
论文解读

OMNI-AVSR: Towards Unified Multimodal Speech Recognition With Large Language Models

语音识别 | 8.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4649 字 阅读 →
论文解读

PerformSinger: Multimodal Singing Voice Synthesis Leveraging Synchronized Lip Cues from Singing Performance Videos

歌唱语音合成 | 4.5/10

 · 更新于 2026-09-25 · 约 3 分钟 · 1471 字 阅读 →
论文解读

Prototype-Guided Cross-Modal Contrastive Learning for Continual Audio-Visual Sound Separation

语音分离 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4252 字 阅读 →
论文解读

PSTalker: Realistic 3D Talking Head Synthesis via a Semantic-Aware Audio-Driven Point-Based Shape

说话人合成 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4433 字 阅读 →
论文解读

Purification Before Fusion: Toward Mask-Free Speech Enhancement for Robust Audio-Visual Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4265 字 阅读 →
论文解读

RAP: Real-Time Audio-Driven Portrait Animation with Video Diffusion Transformer

音视频 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5050 字 阅读 →