论文解读

SNAP-UQ: Self-supervised Next-Activation Prediction for Single-Pass Uncertainty in TinyML

音频分类 | 7.0/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5831 字 阅读 →
论文解读

SpeechJudge: Towards Human-Level Judgment for Speech Naturalness

模型评估 | 8.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5451 字 阅读 →
论文解读

STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence

音频问答 | 8.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4348 字 阅读 →
论文解读

Steering Autoregressive Music Generation with Recursive Feature Machines

音乐生成 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4217 字 阅读 →
论文解读

SyncTrack: Rhythmic Stability and Synchronization in Multi-Track Music Generation

音乐生成 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5626 字 阅读 →
论文解读

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

音频生成 | 8.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4891 字 阅读 →
论文解读

The Deleuzian Representation Hypothesis

模型评估 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4035 字 阅读 →
论文解读

TTSDS2: Resources and Benchmark for Evaluating Human-Quality Text to Speech Systems

模型评估 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4181 字 阅读 →
论文解读

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs

基准测试 | 8.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4207 字 阅读 →
论文解读

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models

基准测试 | 9.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4501 字 阅读 →
论文解读

Audio Video Verbal Analysis (AVVA) for Capturing Classroom Dialogues

音频问答 | 6.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4011 字 阅读 →
论文解读

HATS: An Open data set Integrating Human Perception Applied to the Evaluation of Automatic Speech Recognition Metrics

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3963 字 阅读 →
论文解读

Identifying and typifying demographic unfairness in phoneme-level embeddings of self-supervised speech recognition models

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4560 字 阅读 →
论文解读

MCPHunt: An Evaluation Framework for Cross-Boundary Data Propagation in Multi-Server MCP Agents

模型评估 | 7.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6271 字 阅读 →
论文解读

Qualitative Evaluation of Language Model Rescoring in Automatic Speech Recognition

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 5 分钟 · 2252 字 阅读 →
论文解读

Transformer-Based Rhythm Quantization of Performance MIDI Using Beat Annotations

音乐信息检索 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4287 字 阅读 →
论文解读

A Toolkit for Detecting Spurious Correlations in Speech Datasets

模型评估 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4619 字 阅读 →
论文解读

Preferences of a Voice-First Nation: Large-Scale Pairwise Evaluation and Preference Analysis for TTS in Indian Languages

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4165 字 阅读 →
论文解读

PSP: An Interpretable Per-Dimension Accent Benchmark for Indic Text-to-Speech

语音合成 | 9.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4333 字 阅读 →
论文解读

Tatemae: Detecting Alignment Faking via Tool Selection in LLMs

大语言模型 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4820 字 阅读 →