论文解读

Korean aegyo speech shows systematic F1 increase to signal childlike qualities

语音情感识别 | 6.0/10

 · 更新于 2026-09-10 · 约 5 分钟 · 2463 字 阅读 →
论文解读

Mitigating Shared-Private Branch Imbalance via Dual-Branch Rebalancing for Multimodal Sentiment Analysis

多模态模型 | 7.5/10

 · 更新于 2026-09-10 · 约 11 分钟 · 5421 字 阅读 →
论文解读

ML-SAN: Multi-Level Speaker-Adaptive Network for Emotion Recognition in Conversations

语音情感识别 | 8.0/10

 · 更新于 2026-09-10 · 约 9 分钟 · 4331 字 阅读 →
论文解读

MMEB-V3: Measuring the Performance Gaps of Omni-Modality Embedding Models

基准测试 | 7.5/10

 · 更新于 2026-09-10 · 约 9 分钟 · 4488 字 阅读 →
论文解读

Monitoring exposure-length variations in submarine power cables using distributed fiber-optic sensing

音频事件检测 | 6.5/10

 · 更新于 2026-09-10 · 约 7 分钟 · 3073 字 阅读 →
论文解读

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation

音频生成 | 7.5/10

 · 更新于 2026-09-10 · 约 12 分钟 · 5908 字 阅读 →
论文解读

Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence

多模态模型 | 8.5/10

 · 更新于 2026-09-10 · 约 14 分钟 · 6861 字 阅读 →
论文解读

Praxy Voice: Voice-Prompt Recovery + BUPS for Commercial-Class Indic TTS from a Frozen Non-Indic Base at Zero Commercial-Training-Data Cost

语音合成 | 8.0/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4562 字 阅读 →
论文解读

PSP: An Interpretable Per-Dimension Accent Benchmark for Indic Text-to-Speech

基准测试 | 7.5/10

 · 更新于 2026-09-10 · 约 11 分钟 · 5099 字 阅读 →
论文解读

RAS: a Reliability Oriented Metric for Automatic Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-10 · 约 9 分钟 · 4136 字 阅读 →
论文解读

Robust Accent Identification via Voice Conversion and Non-Timbral Embeddings

语音识别 | 7.5/10

 · 更新于 2026-09-10 · 约 7 分钟 · 3282 字 阅读 →
论文解读

Step-Audio-R1.5 Technical Report

语音对话系统 | 8.0/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4640 字 阅读 →
论文解读

SymphonyGen: 3D Hierarchical Orchestral Generation with Controllable Harmony Skeleton

音乐生成 | 7.5/10

 · 更新于 2026-09-10 · 约 12 分钟 · 5548 字 阅读 →
论文解读

The Structured Output Benchmark: A Multi-Source Benchmark for Evaluating Structured Output Quality in Large Language Models

基准测试 | 7.0/10

 · 更新于 2026-09-10 · 约 8 分钟 · 3637 字 阅读 →
论文解读

UNet-Based Fusion and Exponential Moving Average Adaptation for Noise-Robust Speaker Recognition

说话人验证 | 7.5/10

 · 更新于 2026-09-10 · 约 11 分钟 · 5394 字 阅读 →
论文解读

Unrequited Emotions: Investigating the Gaps in Motivation and Practice in Speech Emotion Recognition Research

语音情感识别 | 8.0/10

 · 更新于 2026-09-10 · 约 6 分钟 · 2880 字 阅读 →
论文解读

Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models

音频问答 | 7.5/10

 · 更新于 2026-09-10 · 约 8 分钟 · 3527 字 阅读 →
论文解读

WhisperPipe: A Resource-Efficient Streaming Architecture for Real-Time Automatic Speech Recognition

语音识别 | 6.5/10

 · 更新于 2026-09-10 · 约 7 分钟 · 3139 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-04-29

共分析 29 篇语音/AI 论文

 · 更新于 2026-09-10 · 约 87 分钟 · 43144 字 阅读 →
论文解读

A Functorial Formulation of Neighborhood Aggregating Deep Learning

理论分析 | 6.5/10

 · 更新于 2026-09-10 · 约 7 分钟 · 3389 字 阅读 →