论文解读

SpeakerCard-1M: An Evidence-Grounded Speaker Card Corpus for In-the-Wild Speaker Verification

说话人验证 | 7.4/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6193 字 阅读 →
论文解读

MOSS-Audio Technical Report

语音识别 | 9.2/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5940 字 阅读 →
论文解读

SN-WER: Script-Normalized WER for Multi-Script Indic ASR Evaluation

语音识别 | 5.3/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6052 字 阅读 →
论文解读

Improving acoustic drone detection generalization through pretraining and data augmentation

音频事件检测 | 7.7/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5246 字 阅读 →
论文解读

Beyond Binary: Speech Representations Across the Cognitive Score Hierarchy

自监督学习 | 8.1/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6023 字 阅读 →
论文解读

FalAR: A Large-scale Speaker-Annotated European Portuguese Speech Corpus of Parliamentary Sessions

语音识别 | 5.5/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5287 字 阅读 →
论文解读

Frame-Aligned Fusion of Canary and WavLM for Non-Intrusive Intelligibility Prediction of Hearing-Aid-Processed Speech

语音质量评估 | 10/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6946 字 阅读 →
论文解读

Speech Quality Embeddings for Improved Detection and Classification of Degradations in Speech Signals

语音质量评估 | 5.8/10

 · 更新于 2026-09-24 · 约 20 分钟 · 9969 字 阅读 →
论文解读

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis

语音合成 | 6.8/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7239 字 阅读 →
论文解读

Beyond Content: A Comprehensive Speech Toxicity Dataset and Detection Framework Incorporating Paralinguistic Cues

音频分类 | 6.5/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7338 字 阅读 →
论文解读

AudioMosaic: Contrastive Masked Audio Representation Learning

音频分类 | 7.3/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7392 字 阅读 →
论文解读

FSD50K-Solo: Automated Curation of Single-Source Sound Events

数据清洗 | 5.5/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6839 字 阅读 →
论文解读

Masked Autoencoders with Limited Data: Does It Work? A Fine-Grained Bioacoustics Case Study

音频分类 | 5.5/10

 · 更新于 2026-09-24 · 约 19 分钟 · 9405 字 阅读 →
论文解读

Text-Dependent Speaker Verification (TdSV) Challenge 2024: Team Naive System Report

说话人验证 | 5.5/10

 · 更新于 2026-09-24 · 约 20 分钟 · 9548 字 阅读 →
论文解读

Bypassing Direct Reconstruction: Speech Detection from MEG via Large-Scale Audio Retrieval

语音活动检测 | 7.0/10

 · 更新于 2026-09-24 · 约 20 分钟 · 9877 字 阅读 →
论文解读

CORTEG: Foundation Models Enable Cross-Modality Representation Transfer from Scalp to Intracranial Brain Recordings

脑机接口 | 6.5/10

 · 更新于 2026-09-24 · 约 19 分钟 · 9166 字 阅读 →
论文解读

Voice Biomarkers for Depression and Anxiety

语音生物标志物 | 1.0/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4782 字 阅读 →
论文解读

Anisotropic Modality Align

Anisotropic Modality Align

 · 更新于 2026-09-24 · 约 16 分钟 · 7980 字 阅读 →
论文解读

BeeVe: Unsupervised Acoustic State Discovery in Honey Bee Buzzing

BeeVe: Unsupervised Acoustic State Discovery in Honey Bee Buzzing

 · 更新于 2026-09-24 · 约 20 分钟 · 9725 字 阅读 →
论文解读

Sparse Autoencoders as Plug-and-Play Firewalls for Adversarial Attack Detection in VLMs

Sparse Autoencoders as Plug-and-Play Firewalls for Adversarial Attack Detection in VLMs

 · 更新于 2026-09-24 · 约 18 分钟 · 8609 字 阅读 →