论文解读

PhoStream: Benchmarking Real-World Streaming for Omnimodal Assistants in Mobile Scenarios

PhoStream: Benchmarking Real-World Streaming for Omnimodal Assistants in Mobile Scenarios

 · 更新于 2026-09-09 · 约 1 分钟 · 36 字 阅读 →
论文解读

Pianist Transformer: Towards Expressive Piano Performance Rendering via Scalable Self-Supervised Pre-Training

Pianist Transformer: Towards Expressive Piano Performance Rendering via Scalable Self-Supervised Pre-Training

 · 更新于 2026-09-09 · 约 1 分钟 · 37 字 阅读 →
论文解读

Polyphonia: Training-Free Context-Aware Music Editing with Acoustic-Informed Attention Calibration

Polyphonia: Training-Free Context-Aware Music Editing with Acoustic-Informed Attention Calibration

 · 更新于 2026-09-09 · 约 1 分钟 · 35 字 阅读 →
论文解读

Position: *Beyond Text* The Text-Centric Bias in Foundation Models Must Be Revisited for a Speech-First Future

Position: *Beyond Text* The Text-Centric Bias in Foundation Models Must Be Revisited for a Speech-First Future

 · 更新于 2026-09-09 · 约 1 分钟 · 42 字 阅读 →
论文解读

Position: Towards Responsible Evaluation for Text-to-Speech

Position: Towards Responsible Evaluation for Text-to-Speech

 · 更新于 2026-09-09 · 约 1 分钟 · 32 字 阅读 →
论文解读

PRIM:Cooperative Dynamic Token Compression for Efficient Large Multimodal Models

PRIM:Cooperative Dynamic Token Compression for Efficient Large Multimodal Models

 · 更新于 2026-09-09 · 约 1 分钟 · 50 字 阅读 →
论文解读

ProactiveLLM: Learning Active Interaction for Streaming Large Language Models

ProactiveLLM: Learning Active Interaction for Streaming Large Language Models

 · 更新于 2026-09-09 · 约 1 分钟 · 35 字 阅读 →
论文解读

Probing Cross-modal Information Hubs in Audio-Visual LLMs

Probing Cross-modal Information Hubs in Audio-Visual LLMs

 · 更新于 2026-09-09 · 约 1 分钟 · 33 字 阅读 →
论文解读

Quaternion Self-Attention with Shared Scores

Quaternion Self-Attention with Shared Scores

 · 更新于 2026-09-09 · 约 1 分钟 · 31 字 阅读 →
论文解读

Query-Based Asymmetric Modeling with Decoupled Input–Output Rates for Speech Restoration

Query-Based Asymmetric Modeling with Decoupled Input–Output Rates for Speech Restoration

 · 更新于 2026-09-09 · 约 1 分钟 · 47 字 阅读 →
论文解读

Real-World Unsupervised Models Generalize to Predict Brain Responses to Out-of-Distribution Stimuli

Real-World Unsupervised Models Generalize to Predict Brain Responses to Out-of-Distribution Stimuli

 · 更新于 2026-09-09 · 约 1 分钟 · 37 字 阅读 →
论文解读

Reasoning LLM Improves Speaker Recognition in Long-form TV Dramas

Reasoning LLM Improves Speaker Recognition in Long-form TV Dramas

 · 更新于 2026-09-09 · 约 1 分钟 · 35 字 阅读 →
论文解读

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models

 · 更新于 2026-09-09 · 约 1 分钟 · 36 字 阅读 →
论文解读

REST: Diffusion-based Real-time End-to-end Streaming Talking Head Generation via ID-Context Caching and Asynchronous Streaming Distillation

REST: Diffusion-based Real-time End-to-end Streaming Talking Head Generation via ID-Context Caching and Asynchronous Streaming Distillation

 · 更新于 2026-09-09 · 约 1 分钟 · 41 字 阅读 →
论文解读

Rethinking Attention in Spiking Transformers: Overcoming Density Bias with Set Similarity

Rethinking Attention in Spiking Transformers: Overcoming Density Bias with Set Similarity

 · 更新于 2026-09-09 · 约 1 分钟 · 37 字 阅读 →
论文解读

Robust Signal Enhancement via Fractional Detail Views and Knowledge Guided Multi-view Fusion

Robust Signal Enhancement via Fractional Detail Views and Knowledge Guided Multi-view Fusion

 · 更新于 2026-09-09 · 约 1 分钟 · 38 字 阅读 →
论文解读

S3Audio: Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer

S3Audio: Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer

 · 更新于 2026-09-09 · 约 1 分钟 · 37 字 阅读 →
论文解读

SALSA-V: Shortcut-Augmented Long-form Synchronized Audio from Videos

SALSA-V: Shortcut-Augmented Long-form Synchronized Audio from Videos

 · 更新于 2026-09-09 · 约 1 分钟 · 33 字 阅读 →
论文解读

SAM Audio: Segment Anything in Audio

** | 6.5/10

 · 更新于 2026-09-09 · 约 1 分钟 · 46 字 阅读 →
论文解读

SARSteer: Safeguarding Large Audio Language Models via Safe-Ablated Refusal Steering

SARSteer: Safeguarding Large Audio Language Models via Safe-Ablated Refusal Steering

 · 更新于 2026-09-09 · 约 1 分钟 · 36 字 阅读 →