论文解读

CTC-DRO: Robust Optimization for Reducing Language Disparities in Speech Recognition

语音识别 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4981 字 阅读 →
论文解读

Deep Learning with Learnable Product-Structured Activations

音频分类 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5033 字 阅读 →
论文解读

DiVeQ: Differentiable Vector Quantization Using the Reparameterization Trick

语音编码 | 8.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6275 字 阅读 →
论文解读

Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception

音频场景理解 | 9.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4566 字 阅读 →
论文解读

OptMerge: Unifying Multimodal LLM Capabilities and Modalities via Model Merging

多模态模型 | 8.0/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6080 字 阅读 →
论文解读

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

音频生成 | 8.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4891 字 阅读 →
论文解读

TINY BUT MIGHTY: A SOFTWARE-HARDWARE CO- DESIGN APPROACH FOR EFFICIENT MULTIMODAL IN- FERENCE ON BATTERY-POWERED SMALL DEVICES

多模态模型 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5357 字 阅读 →
论文解读

Audio Video Verbal Analysis (AVVA) for Capturing Classroom Dialogues

音频问答 | 6.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4011 字 阅读 →
论文解读

MCPHunt: An Evaluation Framework for Cross-Boundary Data Propagation in Multi-Server MCP Agents

模型评估 | 7.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6271 字 阅读 →
论文解读

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis

语音质量评估 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5548 字 阅读 →
论文解读

Praxy Voice: Voice-Prompt Recovery + BUPS for Commercial-Class Indic TTS from a Frozen Non-Indic Base at Zero Commercial-Training-Data Cost

语音合成 | 8.0/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5558 字 阅读 →
论文解读

A Text-To-Text Alignment Algorithm for Better Evaluation of Modern Speech Recognition Systems

模型评估 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3720 字 阅读 →
论文解读

Constructing Composite Features for Interpretable Music-Tagging

音乐信息检索 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4319 字 阅读 →
论文解读

Denoising Of Stochastic Ray Tracing Room Impulse Responses

空间音频 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4886 字 阅读 →
论文解读

ECHO: Frequency-Aware Hierarchical Encoding for Variable-Length Signals

音频分类 | 9.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4374 字 阅读 →
论文解读

Evaluating High-Resolution Piano Sustain Pedal Depth Estimation with Musically Informed Metrics

音乐信息检索 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4473 字 阅读 →
论文解读

MNV-17: A High-Quality Performative Mandarin Dataset for Nonverbal Vocalization Recognition in Speech

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3597 字 阅读 →
论文解读

Polynomial Mixing for Efficient Self-Supervised Speech Encoders

语音识别 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4653 字 阅读 →
论文解读

Praxy Voice: Voice-Prompt Recovery + BUPS for Commercial-Class Indic TTS from a Frozen Non-Indic Base at Zero Commercial-Training-Data Cost

语音合成 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4562 字 阅读 →
论文解读

SA-SSL-MOS: Self-Supervised Learning MOS Prediction with Spectral Augmentation for Generalized Multi-Rate Speech Assessment

语音质量评估 | 7.0/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5780 字 阅读 →