论文解读

Optimality of FSQ Tokens for Continuous Diffusion for Categorical Data with Application to Text-to-Speech

语音合成 | 8/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5305 字 阅读 →
论文解读

PADS-TAL: Padding-Annealed Diffusion Sampling in Text-Aware Latent Space for Robust and Diverse Text-to-Music Generation

音乐生成 | 6.6/10

 · 更新于 2026-09-24 · 约 21 分钟 · 10222 字 阅读 →
论文解读

PCRNet: Phase-aware Complex Refinement Network for EEG-based Auditory Attention Decoding

实时处理 | 6.4/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6788 字 阅读 →
论文解读

PHALAR: Phasors for Learned Musical Audio Representations

音乐生成 | 8/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7969 字 阅读 →
论文解读

PhaseCoder: Microphone Geometry-Agnostic Spatial Audio Understanding for Multimodal LLMs

空间音频 | 8.7/10

 · 更新于 2026-09-24 · 约 17 分钟 · 8299 字 阅读 →
论文解读

PhoStream: Benchmarking Real-World Streaming for Omnimodal Assistants in Mobile Scenarios

音视频问答 | 7.3/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7522 字 阅读 →
论文解读

Pianist Transformer: Towards Expressive Piano Performance Rendering via Scalable Self-Supervised Pre-Training

音乐生成 | 8.1/10

 · 更新于 2026-09-24 · 约 6 分钟 · 2789 字 阅读 →
论文解读

Polyphonia: Zero-Shot Timbre Transfer in Polyphonic Music with Acoustic-Informed Attention Calibration

音乐生成 | 6.5/10

 · 更新于 2026-09-24 · 约 18 分钟 · 8684 字 阅读 →
论文解读

PRIM:Cooperative Dynamic Token Compression for Efficient Large Multimodal Models

音视频理解 | 3.6/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5669 字 阅读 →
论文解读

ProactiveLLM: Learning Active Interaction for Streaming Large Language Models

语音识别 | 7.2/10

 · 更新于 2026-09-24 · 约 8 分钟 · 3736 字 阅读 →
论文解读

Probing Cross-modal Information Hubs in Audio-Visual LLMs

音视频理解 | 7.2/10

 · 更新于 2026-09-24 · 约 6 分钟 · 2684 字 阅读 →
论文解读

Quaternion Self-Attention with Shared Scores

语音增强 | 6.3/10

 · 更新于 2026-09-24 · 约 21 分钟 · 10452 字 阅读 →
论文解读

Query-Based Asymmetric Modeling with Decoupled Input–Output Rates for Speech Restoration

语音增强 | 7.1/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6737 字 阅读 →
论文解读

Real-World Unsupervised Models Generalize to Predict Brain Responses to Out-of-Distribution Stimuli

模型评估 | 6.9/10

 · 更新于 2026-09-24 · 约 20 分钟 · 9530 字 阅读 →
论文解读

Reasoning LLM Improves Speaker Recognition in Long-form TV Dramas

Reasoning LLM Improves Speaker Recognition in Long-form TV Dramas

 · 更新于 2026-09-24 · 约 15 分钟 · 7136 字 阅读 →
论文解读

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models

语音编码 | 8/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6645 字 阅读 →
论文解读

REST: Diffusion-based Real-time End-to-end Streaming Talking Head Generation via ID-Context Caching and Asynchronous Streaming Distillation

音视频生成 | 7.3/10

 · 更新于 2026-09-24 · 约 26 分钟 · 12601 字 阅读 →
论文解读

Rethinking Attention in Spiking Transformers: Overcoming Density Bias with Set Similarity

音频分类 | 3.6/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7033 字 阅读 →
论文解读

Robust Signal Enhancement via Fractional Detail Views and Knowledge Guided Multi-view Fusion

语音增强 | 5.7/10

 · 更新于 2026-09-24 · 约 17 分钟 · 8257 字 阅读 →
论文解读

SALSA-V: Shortcut-Augmented Long-form Synchronized Audio from Videos

音视频生成 | 7.6/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7734 字 阅读 →