论文解读
JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments
声源定位 | 8.1/10
声源定位 | 8.1/10
语音识别 | 9.3/10
语音交互 | 8.1/10
语音属性识别 | 5.4/10
语音伪造检测 | 9.3/10
模型压缩 | 6.3/10
音频修复 | 7.4/10
音频理解 | 9.1/10
音频理解 | 6.4/10
音视频理解 | 5.4/10
音频分类 | 6.5/10
MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models
MoST: Mixing Speech and Text with Modality-Aware Mixture of Experts
音频理解 | 6.4/10
鲁棒性 | 5.5/10
语音合成 | 6.1/10
多模态模型 | 5.2/10
多模态模型 | 8/10
音频伪造检测 | 6.1/10
音频事件检测 | 5.5/10