论文解读

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents

基准测试 | 7.0/10

 · 更新于 2026-10-02 · 约 8 分钟 · 3824 字 阅读 →
论文解读

Cross-Linguistic Rhythmic and Spectral Feature-Based Analysis of Nyishi and Adi: Two Under-Resourced Languages of Arunachal Pradesh

Cross-Linguistic Rhythmic and Spectral Feature-Based Analysis of Nyishi and Adi: Two Under-Resourced Languages of Arunachal Pradesh

 · 更新于 2026-10-02 · 约 1 分钟 · 34 字 阅读 →
论文解读

Cutscene Agent: An LLM Agent Framework for Automated 3D Cutscene Generation

生成模型 | 8.5/10

 · 更新于 2026-10-02 · 约 12 分钟 · 5619 字 阅读 →
论文解读

缺失与偏置并存时如何做鲁棒的多模态情感分析:门控序列修复与平衡跨模态注意的协同

📄 缺失与偏置并存时如何做鲁棒的多模态情感分析:门控序列修复与平衡跨模态注意的协同 会议论文 ID:conference:icassp:2026:icassp-arnumber:11460389

 · 更新于 2026-10-02 · 约 16 分钟 · 7690 字 阅读 →
论文解读

Generative UI as an Accessibility Bridge: Lessons from C2C E-Commerce

无障碍 | 6.5/10

 · 更新于 2026-10-02 · 约 9 分钟 · 4025 字 阅读 →
论文解读

Huí Sù: Co-constructing a Dual Feedback Apparatus

音乐生成 | 5.5/10

 · 更新于 2026-10-02 · 约 7 分钟 · 3319 字 阅读 →
论文解读

Human-1 by Josh Talks: A Full-Duplex Conversational Modeling Framework in Hindi using Real-World Conversations

语音对话系统 | 7.5/10

 · 更新于 2026-10-02 · 约 12 分钟 · 5611 字 阅读 →
论文解读

Independent-Component-Based Encoding Models of Brain Activity During Story Comprehension

神经编码 | 7.5/10

 · 更新于 2026-10-02 · 约 12 分钟 · 5890 字 阅读 →
论文解读

保留高维 XLS-R 特征时单层 KAN 能否以更少参数替代复杂后端

📄 保留高维 XLS-R 特征时单层 KAN 能否以更少参数替代复杂后端 会议论文 ID:conference:icassp:2026:icassp-arnumber:11460320

 · 更新于 2026-10-02 · 约 18 分钟 · 8894 字 阅读 →
论文解读

Korean aegyo speech shows systematic F1 increase to signal childlike qualities

语音情感识别 | 6.0/10

 · 更新于 2026-10-02 · 约 5 分钟 · 2463 字 阅读 →
论文解读

Mitigating Shared-Private Branch Imbalance via Dual-Branch Rebalancing for Multimodal Sentiment Analysis

多模态模型 | 7.5/10

 · 更新于 2026-10-02 · 约 11 分钟 · 5421 字 阅读 →
论文解读

用带噪混合当高噪声步监督:Mix2Morph 如何把加法叠加训练成可控的声音注入

📄 用带噪混合当高噪声步监督:Mix2Morph 如何把加法叠加训练成可控的声音注入 会议论文 ID:conference:icassp:2026:icassp-arnumber:11460386

 · 更新于 2026-10-02 · 约 16 分钟 · 7807 字 阅读 →
论文解读

ML-SAN: Multi-Level Speaker-Adaptive Network for Emotion Recognition in Conversations

语音情感识别 | 8.0/10

 · 更新于 2026-10-02 · 约 9 分钟 · 4331 字 阅读 →
论文解读

MMEB-V3: Measuring the Performance Gaps of Omni-Modality Embedding Models

基准测试 | 7.5/10

 · 更新于 2026-10-02 · 约 9 分钟 · 4488 字 阅读 →
论文解读

Monitoring exposure-length variations in submarine power cables using distributed fiber-optic sensing

音频事件检测 | 6.5/10

 · 更新于 2026-10-02 · 约 7 分钟 · 3073 字 阅读 →
论文解读

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation

音频生成 | 7.5/10

 · 更新于 2026-10-02 · 约 12 分钟 · 5908 字 阅读 →
论文解读

Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence

多模态模型 | 8.5/10

 · 更新于 2026-10-02 · 约 14 分钟 · 6861 字 阅读 →
论文解读

Praxy Voice: Voice-Prompt Recovery + BUPS for Commercial-Class Indic TTS from a Frozen Non-Indic Base at Zero Commercial-Training-Data Cost

语音合成 | 8.0/10

 · 更新于 2026-10-02 · 约 10 分钟 · 4562 字 阅读 →
论文解读

PSP: An Interpretable Per-Dimension Accent Benchmark for Indic Text-to-Speech

基准测试 | 7.5/10

 · 更新于 2026-10-02 · 约 11 分钟 · 5099 字 阅读 →
论文解读

RAS: a Reliability Oriented Metric for Automatic Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-10-02 · 约 9 分钟 · 4136 字 阅读 →