论文解读

Robust Accent Identification via Voice Conversion and Non-Timbral Embeddings

语音识别 | 7.5/10

 · 更新于 2026-10-02 · 约 7 分钟 · 3282 字 阅读 →
论文解读

Step-Audio-R1.5 Technical Report

语音对话系统 | 8.0/10

 · 更新于 2026-10-02 · 约 10 分钟 · 4640 字 阅读 →
论文解读

SymphonyGen: 3D Hierarchical Orchestral Generation with Controllable Harmony Skeleton

音乐生成 | 7.5/10

 · 更新于 2026-10-02 · 约 12 分钟 · 5548 字 阅读 →
论文解读

The Structured Output Benchmark: A Multi-Source Benchmark for Evaluating Structured Output Quality in Large Language Models

基准测试 | 7.0/10

 · 更新于 2026-10-02 · 约 8 分钟 · 3637 字 阅读 →
论文解读

UNet-Based Fusion and Exponential Moving Average Adaptation for Noise-Robust Speaker Recognition

说话人验证 | 7.5/10

 · 更新于 2026-10-02 · 约 11 分钟 · 5394 字 阅读 →
论文解读

Unrequited Emotions: Investigating the Gaps in Motivation and Practice in Speech Emotion Recognition Research

语音情感识别 | 8.0/10

 · 更新于 2026-10-02 · 约 6 分钟 · 2880 字 阅读 →
论文解读

Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models

音频问答 | 7.5/10

 · 更新于 2026-10-02 · 约 8 分钟 · 3527 字 阅读 →
论文解读

WhisperPipe: A Resource-Efficient Streaming Architecture for Real-Time Automatic Speech Recognition

语音识别 | 6.5/10

 · 更新于 2026-10-02 · 约 7 分钟 · 3139 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-04-29

共分析 29 篇语音/AI 论文

 · 更新于 2026-10-02 · 约 87 分钟 · 43144 字 阅读 →
论文解读

A Functorial Formulation of Neighborhood Aggregating Deep Learning

理论分析 | 6.5/10

 · 更新于 2026-10-02 · 约 7 分钟 · 3389 字 阅读 →
论文解读

All That Glitters Is Not Audio: Rethinking Text Priors and Audio Reliance in Audio-Language Evaluation

音频问答 | 6.5/10

 · 更新于 2026-10-02 · 约 9 分钟 · 4036 字 阅读 →
论文解读

An event-based sequence modeling approach to recognizing non-triad chords with oversegmentation minimization

音乐理解 | 7.5/10

 · 更新于 2026-10-02 · 约 9 分钟 · 4400 字 阅读 →
论文解读

CineAGI: Character-Consistent Movie Creation through LLM-Orchestrated Multi-Modal Generation and Cross-Scene Integration

跨模态 | 8.0/10

 · 更新于 2026-10-02 · 约 9 分钟 · 4377 字 阅读 →
论文解读

Come Together: Analyzing Popular Songs Through Statistical Embeddings

音乐信息检索 | 6.5/10

 · 更新于 2026-10-02 · 约 8 分钟 · 3526 字 阅读 →
论文解读

Comparison of sEMG Encoding Accuracy Across Speech Modes Using Articulatory and Phoneme Features

语音生物标志物 | 8.0/10

 · 更新于 2026-10-02 · 约 9 分钟 · 4365 字 阅读 →
论文解读

Explainable AI in Speaker Recognition -- Making Latent Representations Understandable

说话人识别 | 7.5/10

 · 更新于 2026-10-02 · 约 8 分钟 · 3672 字 阅读 →
论文解读

Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation

音视频 | 8.5/10

 · 更新于 2026-10-02 · 约 12 分钟 · 5626 字 阅读 →
论文解读

HeadRouter: Dynamic Head-Weight Routing for Task-Adaptive Audio Token Pruning in Large Audio Language Models

音频大模型 | 8.0/10

 · 更新于 2026-10-02 · 约 11 分钟 · 5044 字 阅读 →
论文解读

Latent-Hysteresis Graph ODEs: Modeling Coupled Topology-Feature Evolution via Continuous Phase Transitions

图神经网络 | 8.0/10

 · 更新于 2026-10-02 · 约 10 分钟 · 4777 字 阅读 →
论文解读

Listening with Time: Precise Temporal Awareness for Long-Form Audio Understanding

音频场景理解 | 8.0/10

 · 更新于 2026-10-02 · 约 10 分钟 · 4909 字 阅读 →