论文解读

A Dataset and Benchmark for Optical Music Recognition of String Quartet Scores

端到端 | 8.5/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7765 字 阅读 →
论文解读

DIY e-HandPan: A new DIY Low-Cost Handpan Interface based on Arduino and ESP32 Microcontrollers

音频交互 | 7.2/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6288 字 阅读 →
论文解读

DuplexWorld: Can voice agents help you get through the day?

语音交互 | 8.5/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8820 字 阅读 →
论文解读

Monophonic Audio Synthesizer Using FPGAs

音频生成 | 4.2/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5225 字 阅读 →
论文解读

TimeRoute: Time-Aware Modality Routing and Diffusion for Multi-Modal Recommendation

扩散模型 | 6.5/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9515 字 阅读 →
论文解读

Training Set Synthesis for Bioacoustic Denoising: A Case Study With Mice

语音增强 | 8.0/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8643 字 阅读 →
论文解读

BAMU: Bitstream-Aware Marginal-Utility Allocation for Frozen Pretrained Neural Speech Codecs

语音编码 | 7.1/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6455 字 阅读 →
论文解读

Investigating Multimodal Informativity under Different Partner Visibility Conditions in Video-Mediated Dialogue

多模态模型 | 5.9/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7905 字 阅读 →
论文解读

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure

语音编码 | 7.7/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9038 字 阅读 →
论文解读

How to Recognize New Words: A Comparison Between Context Biasing Methods and Speech LLMs

语音识别 | 6.4/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7408 字 阅读 →
论文解读

PD-GS: Phoneme-Driven 3DGS for Audio-Driven Talking Heads

音视频生成 | 5.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5522 字 阅读 →
论文解读

Breaking the Curse of Multilinguality in Many-to-Many Speech-to-Text Translation via a Resource-Aware Mixture of Speech Encoders

语音翻译 | 8.2/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7762 字 阅读 →
论文解读

Crowdsourced Multilingual Speech Intelligibility Testing

语音质量评估 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5390 字 阅读 →
论文解读

Leveraging Beam Search Information for Confidence Estimation in E2E ASR

语音识别 | 6.3/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7632 字 阅读 →
论文解读

Unfolded Recursive Expectation-Maximization Neural Network For Speaker Tracking

声源定位 | 5.4/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5982 字 阅读 →
论文解读

Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models

语音交互 | 7.6/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8635 字 阅读 →
论文解读

End-to-End Markov State Sequence Learning for Auditory Attention Decoding

语音交互 | 8.3/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7112 字 阅读 →
论文解读

What the Waveform Knows: Transparent-first Speech and Audio Intelligence with Caption Studio

语音识别 | 4.8/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6414 字 阅读 →
论文解读

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications

音视频生成 | 7.5/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7604 字 阅读 →
论文解读

SSTMark: Robust Training-Free Semantic-Level Speech Watermarking

音频水印 | 6.5/10

 · 更新于 2026-09-25 · 约 21 分钟 · 10235 字 阅读 →