论文解读

Ultra-Compact CNN Architectures for Tropical Bird Audio Detection on Microcontrollers

音频事件检测 | 9.3/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7444 字 阅读 →
论文解读

Validating the Single Item Kawaii Measure

音频理解 | 6.4/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8270 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-07-23

共分析 21 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 70 分钟 · 34984 字 阅读 →
论文解读

A Situational Speech Synthesizer for Yoruba: System Design, Phonological Rule Architecture, and Orthographic Extensions for Contour

语音合成 | 6.7/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6602 字 阅读 →
论文解读

Addressing Limited Data in Auditory Attention Decoding with Diffusion Generative Models

语音分离 | 5.1/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5908 字 阅读 →
论文解读

Benchmarking Human and Automatic Speech Recognition of Diverse Speech: Initial Results

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8304 字 阅读 →
论文解读

Comparing Spectrogram Front-Ends for Abnormal Heart-Sound Detection with a Convolutional Neural Network

音频分类 | 5.7/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6412 字 阅读 →
论文解读

Constrained CTC Decoding for Efficient Diacritic Restoration

语音识别 | 7.7/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6374 字 阅读 →
论文解读

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances

语音编码 | 9.2/10

 · 更新于 2026-09-06 · 约 18 分钟 · 8587 字 阅读 →
论文解读

CS-ETS: Chaos-Inspired Samba-Based EMG-To-Speech Synthesis with Nonlinear Chaotic Losses

语音合成 | 5.3/10

 · 更新于 2026-09-06 · 约 18 分钟 · 8594 字 阅读 →
论文解读

EmoEUS: Uncertainty Supervision for Multimodal Emotion Recognition in Conversation

语音情感识别 | 5.6/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6251 字 阅读 →
论文解读

End-to-End Markov State Sequence Learning for Auditory Attention Decoding

语音交互 | 8.3/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7112 字 阅读 →
论文解读

Fretiq: Browser-Native Electric Guitar String Classification via Engineered Spectral Features and Held-Out Free-Play Evaluation

音频分类 | 7.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5587 字 阅读 →
论文解读

From a Multilingual Streaming ASR Backbone to Kenyan-Language Systems: Data-Centric Adaptation of Nemotron 3.5 for Kikuyu, Dholuo, and Kalenjin

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4601 字 阅读 →
论文解读

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio

音频检索 | 8.6/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8119 字 阅读 →
论文解读

MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings

基准测试 | 7.2/10

 · 更新于 2026-09-06 · 约 20 分钟 · 9969 字 阅读 →
论文解读

Staged Depth-Pruning Distillation of a Flow-Matching Text-to-Speech Teacher: A Compact Hindi Speech Synthesizer

语音合成 | 7.9/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8461 字 阅读 →
论文解读

Summary of DCASE 2026 Task 5: Audio-Dependent Question Answering

音频理解 | 5.4/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7727 字 阅读 →
论文解读

Teleportation Game: Quantum Teleportation in Multi-Agent Systems for Interactive Music

音乐生成 | 4.4/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6163 字 阅读 →
论文解读

Towards a reproducible cross-venue method for quantifying crowd noise in stadiums

音频质量评估 | 5.4/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6934 字 阅读 →