论文解读

MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction

语音对话系统 | 8.5/10

 · 更新于 2026-10-02 · 约 14 分钟 · 6664 字 阅读 →
论文解读

Normativity and Productivism: Ableist Intelligence? A Degrowth Analysis of AI Sign Language Translation Tools for Deaf People

语音翻译 | 3.5/10

 · 更新于 2026-10-02 · 约 5 分钟 · 2038 字 阅读 →
论文解读

Predicting Upcoming Stuttering Events from Three-Second Audio: Stratified Evaluation Reveals Severity-Selective Precursors, and the Model Deploys Fully On-Device

语音生物标志物 | 7.0/10

 · 更新于 2026-10-02 · 约 13 分钟 · 6353 字 阅读 →
论文解读

Qualitative Evaluation of Language Model Rescoring in Automatic Speech Recognition

语音识别 | 6.5/10

 · 更新于 2026-10-02 · 约 5 分钟 · 2252 字 阅读 →
论文解读

Selective Augmentation: Improving Universal Automatic Phonetic Transcription via G2P Bootstrapping

语音识别 | 6.0/10

 · 更新于 2026-10-02 · 约 7 分钟 · 3160 字 阅读 →
论文解读

Spectrographic Portamento Gradient Analysis: A Quantitative Method for Historical Cello Recordings with Application to Beethoven's Piano and Cello Sonatas, 1930--2012

音乐信息检索 | 7.5/10

 · 更新于 2026-10-02 · 约 7 分钟 · 3274 字 阅读 →
论文解读

Taming Noise-Induced Prototype Degradation for Privacy-Preserving Personalized Federated Fine-Tuning

个性化联邦学习 | 7.5/10

 · 更新于 2026-10-02 · 约 7 分钟 · 3077 字 阅读 →
论文解读

Transformer-Based Rhythm Quantization of Performance MIDI Using Beat Annotations

音乐信息检索 | 8.0/10

 · 更新于 2026-10-02 · 约 9 分钟 · 4287 字 阅读 →
论文解读

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis

语音质量评估 | 7.5/10

 · 更新于 2026-10-02 · 约 12 分钟 · 5548 字 阅读 →
论文解读

UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions

音频生成 | 8.5/10

 · 更新于 2026-10-02 · 约 12 分钟 · 5965 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-01

共分析 21 篇语音/AI 论文

 · 更新于 2026-10-02 · 约 70 分钟 · 35064 字 阅读 →
论文解读

A New Location Estimator for Mixed LOS & NLOS scenarios

声源定位 | 7.5/10

 · 更新于 2026-10-02 · 约 12 分钟 · 5838 字 阅读 →
论文解读

A Toolkit for Detecting Spurious Correlations in Speech Datasets

模型评估 | 7.0/10

 · 更新于 2026-10-02 · 约 10 分钟 · 4619 字 阅读 →
论文解读

DiffAnon: Diffusion-based Prosody Control for Voice Anonymization

语音匿名化 | 7.5/10

 · 更新于 2026-10-02 · 约 10 分钟 · 4816 字 阅读 →
论文解读

Diffusion Reconstruction towards Generalizable Audio Deepfake Detection

音频深度伪造检测 | 7.5/10

 · 更新于 2026-10-02 · 约 9 分钟 · 4407 字 阅读 →
论文解读

Dual-LoRA: Parameter-Efficient Adversarial Disentanglement for Cross-Lingual Speaker Verification

说话人验证 | 7.5/10

 · 更新于 2026-10-02 · 约 13 分钟 · 6131 字 阅读 →
论文解读

EmoTransCap: Dataset and Pipeline for Emotion Transition-Aware Speech Captioning in Discourses

语音情感识别 | 7.5/10

 · 更新于 2026-10-02 · 约 11 分钟 · 5250 字 阅读 →
论文解读

Fitting Large Nonlinear Mixed Effects Models Using Variational Expectation Maximization

统计计算 | 6.5/10

 · 更新于 2026-10-02 · 约 6 分钟 · 2744 字 阅读 →
论文解读

Full band denoising of room impulse response in the wavelet domain with dictionary learning

音频信号处理 | 6.5/10

 · 更新于 2026-10-02 · 约 10 分钟 · 4520 字 阅读 →
论文解读

Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation

音视频 | 8.5/10

 · 更新于 2026-10-02 · 约 11 分钟 · 5319 字 阅读 →