论文解读

Spectro-Temporal Interference Confounds Phase Encoding in Spatial Audio Foundation Models

自监督学习 | 7.4/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5251 字 阅读 →
论文解读

Frozen Multimodal Embeddings for Personality and Cognitive Ability Assessment in Asynchronous Video Interviews

语音情感识别 | 6.7/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6098 字 阅读 →
论文解读

Assessing the Energy and Carbon Emissions of Neural Speaker Verification Model in Training and Inference

说话人验证 | 7.4/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5733 字 阅读 →
论文解读

OpenBibleTTS: Large-Scale Speech Resources and TTS Models for Low-Resource Languages

语音合成 | 8/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6802 字 阅读 →
论文解读

Automatic Labelling of Speech Translation Errors

语音识别 | 6.1/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4750 字 阅读 →
论文解读

SHALA-LLM: Smartly Handling Ambiguous Labels in Aligning LLMs

语音情感识别 | 6.8/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5860 字 阅读 →
论文解读

OmniHalluc-L: Counterfactual Benchmarking and Modality-Perturbation Reliability Calibration for Long-Form Omni Hallucination

多模态模型 | 7.8/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6119 字 阅读 →
论文解读

Context-aware child-directed speech detection from long-form recordings

自监督学习 | 8.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5403 字 阅读 →
论文解读

Dasheng AudioGen: A Unified Model for Generating Coherent Audio Scenes from Text

音频生成 | 8.6/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7601 字 阅读 →
论文解读

VoiceGiraffe: A Benchmark for Extreme Long-Context Audio-Language Understanding

音频问答 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5379 字 阅读 →
论文解读

Cost-Effective Model Evaluation with Meta-Learning

迁移学习 | 5.4/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7948 字 阅读 →
论文解读

EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation

音视频 | 7.1/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6541 字 阅读 →
论文解读

StepAudio 2.5 Technical Report

统一音频模型 | 8.3/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6187 字 阅读 →
论文解读

UniSRM: A Unified Speech Reward Model for Reasoning-Based Fine-grained Assessment

语音质量评估 | 10/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5850 字 阅读 →
论文解读

Codec-Robust Attacks on Audio LLMs

音频安全 | 8.3/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8462 字 阅读 →
论文解读

CounterFlow: A Two-Phase Inference-Time Sampling for Counterfactual Video Foley Generation

音频生成 | 8.7/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7201 字 阅读 →
论文解读

Evaluating Speech Articulation Synthesis with Articulatory Phoneme Recognition

语音质量评估 | 8.2/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7861 字 阅读 →
论文解读

From Numbers to Perception, Energy Decay Curves Prediction

空间音频 | 7.2/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8331 字 阅读 →
论文解读

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation

基准测试 | 8.1/10

 · 更新于 2026-09-25 · 约 20 分钟 · 9524 字 阅读 →
论文解读

SEABAD: A Tropical Bird Activity Detection Dataset for Passive Acoustic Monitoring

生物声学 音频事件检测 | 8.1/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8250 字 阅读 →