论文解读

Mitigating Spectral Bias in Neural Operators for Underwater Transmission Loss Prediction

音频理解 | 7.3/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4170 字 阅读 →
论文解读

Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities

音频理解 | 7.5/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4569 字 阅读 →
论文解读

SpeechSense: A Paralinguistic-Focused Dataset for Fine-Grained Speech Sentiment Analysis

语音情感识别 | 8.5/10

 · 更新于 2026-09-24 · 约 7 分钟 · 3361 字 阅读 →
论文解读

Uncertainty-Aware Decision Making in Multimodal Large Language Models

音频理解 | 6.8/10

 · 更新于 2026-09-24 · 约 8 分钟 · 3533 字 阅读 →
论文解读

ARENA: Automated Red-Teaming for Large Audio Language Models

音频理解 | 6.5/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7088 字 阅读 →
论文解读

Impulse Response Estimation via Laguerre-Fourier Expansion

音频理解 | 6.8/10

 · 更新于 2026-09-24 · 约 19 分钟 · 9364 字 阅读 →
论文解读

Listen, Reason, and Segment: Aligning LALMs with Editorial Judgment for Media Chapterization

音频理解 | 6.8/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7525 字 阅读 →
论文解读

Leading-Silence Augmentation and Multi-Stage Synthetic Supervision for the Second MLC-SLM Challenge

说话人日志 | 6.5/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6807 字 阅读 →
论文解读

Measuring Fairness in Large Audio Language Models via Semantic-Aware Bias Estimation

语音识别 | 5.6/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7337 字 阅读 →
论文解读

The SLT 2026 SmartGlasses Challenge: Benchmarking Egocentric Multi-Talker Speech Recognition and Understanding with Audio-Language Models

语音识别 | 7.5/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6748 字 阅读 →
论文解读

Modeling and Interpreting Correlations, Null Distributions and Significance Levels in Neural Tracking of Natural Stimuli

音频理解 | 8.0/10

 · 更新于 2026-09-24 · 约 18 分钟 · 8620 字 阅读 →
论文解读

From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs

音频理解 | 7.4/10

 · 更新于 2026-09-24 · 约 5 分钟 · 2056 字 阅读 →
论文解读

VoxZip: Semantic-Anchored Temporal KV Cache Compression for Long-Context Audio Inference

音频理解 | 7.7/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6697 字 阅读 →
论文解读

A Study of ASR Adaptation and Representation Dimensionality Reduction in Persian Speech Emotion Recognition Using Whisper

语音情感识别 | 5.7/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4924 字 阅读 →
论文解读

AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks

语音伪造检测 | 8.4/10

 · 更新于 2026-09-24 · 约 24 分钟 · 11604 字 阅读 →
论文解读

Beyond Residual Connections: Manifold-Constrained Hyper-Connections for Robust Speaker Representation Learning

说话人验证 | 6.0/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4473 字 阅读 →
论文解读

Decolonizing Linguistic Policies in Automated Speech Recognition: A Framework for Cross-Culturally Competent Speech AI

语音识别 | 7.8/10

 · 更新于 2026-09-24 · 约 7 分钟 · 3465 字 阅读 →
论文解读

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model

音乐生成 | 6.5/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6469 字 阅读 →
论文解读

ECHO: A Locally-Deployable Agentic Health Assistant with Temporal Memory, Safety Guardrails, and Speech Assessment

语音属性识别 | 5.5/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7091 字 阅读 →
论文解读

EG-VAE: A Unified Framework for Electric Guitar Tone Transfer and Removal

音频生成 | 7.3/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7237 字 阅读 →