论文解读

Pay Attention to CTC: Fast and Robust Pseudo-Labelling for Unified Speech Recognition

语音识别 | 8.0/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5713 字 阅读 →
论文解读

A Consistent Learning Depression Detection Framework Integrating Multi-View Attention

语音生物标志物 | 6.5/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4619 字 阅读 →
论文解读

A Distribution Matching Approach to Neural Piano Transcription with Optimal Transport

音乐转录 | 7.0/10

 · 更新于 2026-09-24 · 约 8 分钟 · 3781 字 阅读 →
论文解读

Adversarial Rivalry Learning for Music Classification

音乐分类 | 6.5/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4912 字 阅读 →
论文解读

An Audio-Visual Speech Separation Network with Joint Cross-Attention and Iterative Modeling

语音分离 | 7.5/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4638 字 阅读 →
论文解读

Attentive AV-Fusionnet: Audio-Visual Quality Prediction with Hybrid Attention

音视频 | 7.0/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4767 字 阅读 →
论文解读

Caption and Audio-Guided Video Representation Learning with Gated Attention for Partially Relevant Video Retrieval

视频检索 | 7.0/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4469 字 阅读 →
论文解读

Chunk-Wise Attention Transducers for Fast and Accurate Streaming Speech-to-Text

语音识别 | 7.5/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4722 字 阅读 →
论文解读

Combining Multi-Order Attention and Multi-Resolution Discriminator for High-Fidelity Neural Vocoder

语音合成 | 6.5/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6236 字 阅读 →
论文解读

DAT-CFTNet: Speech Enhancement for Cochlear Implant Recipients using Attention-based Dual-Path Recurrent Neural Network

语音增强 | 7.0/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4653 字 阅读 →
论文解读

Distilling Attention Knowledge for Speaker Verification

说话人验证 | 8.0/10

 · 更新于 2026-09-24 · 约 8 分钟 · 3961 字 阅读 →
论文解读

缺失与偏置并存时如何做鲁棒的多模态情感分析:门控序列修复与平衡跨模态注意的协同

📄 缺失与偏置并存时如何做鲁棒的多模态情感分析:门控序列修复与平衡跨模态注意的协同 会议论文 ID:conference:icassp:2026:icassp-arnumber:11460389

 · 更新于 2026-09-24 · 约 16 分钟 · 7690 字 阅读 →
论文解读

Expressive Voice Conversion with Controllable Emotional Intensity

语音转换 | 7.5/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5587 字 阅读 →
论文解读

FDCNet: Frequency Domain Channel Attention and Convolution for Lipreading

视觉语音识别 | 8.5/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4228 字 阅读 →
论文解读

HarmoNet: Music Grounding by Short Video via Harmonic Resample and Dynamic Sparse Alignment

音乐检索 | 7.0/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4256 字 阅读 →
论文解读

Learning What to Hear: Boosting Sound-Source Association for Robust Audiovisual Instance Segmentation

音视频实例分割 | 7.5/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4721 字 阅读 →
论文解读

MFF-RVRDI: Multimodal Fusion Framework for Robust Video Recording Device Identification

视频设备识别 | 7.5/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4466 字 阅读 →
论文解读

MSCT: Differential Cross-Modal Attention for Deepfake Detection

音频深度伪造检测 | 6.5/10

 · 更新于 2026-09-24 · 约 8 分钟 · 3968 字 阅读 →
论文解读

Musicdetr: A Position-Aware Spectral Note Detection Model for Singing Transcription

歌唱语音转录 | 8.5/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4470 字 阅读 →
论文解读

QFOCUS: Controllable Synthesis for Automated Speech Stress Editing to Deliver Human-Like Emphatic Intent

语音合成 | 7.5/10

 · 更新于 2026-09-24 · 约 4 分钟 · 1766 字 阅读 →