论文解读

TINY BUT MIGHTY: A SOFTWARE-HARDWARE CO- DESIGN APPROACH FOR EFFICIENT MULTIMODAL IN- FERENCE ON BATTERY-POWERED SMALL DEVICES

多模态模型 | 7.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5357 字 阅读 →
论文解读

Token-Based Audio Inpainting via Discrete Diffusion

音频生成 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4973 字 阅读 →
论文解读

Toward Complex-Valued Neural Networks for Waveform Generation

语音合成 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4635 字 阅读 →
论文解读

Towards True Speech-to-Speech Models Without Text Guidance

语音对话系统 | 9.1/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4831 字 阅读 →
论文解读

TRIBE: TRImodal Brain Encoder for whole-brain fMRI response prediction

脑编码 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5372 字 阅读 →
论文解读

TripleSumm: Adaptive Triple-Modality Fusion for Video Summarization

视频摘要 | 8.0/10

 · 更新于 2026-09-06 · 约 8 分钟 · 4001 字 阅读 →
论文解读

TTSDS2: Resources and Benchmark for Evaluating Human-Quality Text to Speech Systems

模型评估 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4181 字 阅读 →
论文解读

TVTSyn: Content-Synchronous Time-Varying Timbre for Streaming Voice Conversion and Anonymization

语音转换 | 7.5/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6178 字 阅读 →
论文解读

UALM: Unified Audio Language Model for Understanding, Generation and Reasoning

音频生成 | 8.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4691 字 阅读 →
论文解读

Unified Multi-Modal Interactive and Reactive 3D Motion Generation via Rectified Flow

音频生成 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4970 字 阅读 →
论文解读

UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice

语音翻译 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5273 字 阅读 →
论文解读

Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio Classification

音频分类 | 9.0/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5834 字 阅读 →
论文解读

VibeVoice: Expressive Podcast Generation with Next-Token Diffusion

语音合成 | 8.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5927 字 阅读 →
论文解读

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Video

多模态模型 | 7.0/10

 · 更新于 2026-09-06 · 约 7 分钟 · 3342 字 阅读 →
论文解读

VowelPrompt: Hearing Speech Emotions from Text via Vowel-level Prosodic Augmentation

语音情感识别 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5344 字 阅读 →
论文解读

VoxPrivacy: A Benchmark for Evaluating Interactional Privacy of Speech Language Models

语音对话系统 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4340 字 阅读 →
论文解读

WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM

音频检索 视频检索 | 8.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5230 字 阅读 →
论文解读

WearVox: An Egocentric Multichannel Voice Assistant Benchmark for Wearables

语音对话系统 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5209 字 阅读 →
论文解读

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs

基准测试 | 8.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4207 字 阅读 →
论文解读

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models

基准测试 | 9.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4501 字 阅读 →