论文解读

From Sinhala to Dhivehi: Cross-Lingual Transfer Learning for Low-Resource Speech Recognition

语音识别 | 6.6/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7890 字 阅读 →
论文解读

Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 7010 字 阅读 →
论文解读

Jointly Improving Dialect Identification and ASR in Indian Languages using Multimodal Feature Fusion

语音识别 | 5.8/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7082 字 阅读 →
论文解读

Listen, Think, Transcribe: Continuous Latent Test-Time Scaling for ASR

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6575 字 阅读 →
论文解读

MOSAIC: Interpretable Multi-Token Cross-Attention of Biophonetic and Self-Supervised Representations for Unified Voice Anti-Spoofing

语音伪造检测 | 6.3/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7094 字 阅读 →
论文解读

RABBiT: Rapidly adaptive BOLD foundation model via brain-tuning for accurate zero-shot and few-shot prediction of speech-elicited responses in the brain

音频理解 | 8.1/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6569 字 阅读 →
论文解读

Speaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization

语音编码 | 7.9/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5637 字 阅读 →
论文解读

Towards Language-Agnostic Speech Inversion

语音属性识别 | 5.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5655 字 阅读 →
论文解读

Alethia: a Foundational Encoder for Voice Deepfakes

语音伪造检测 | 7.6/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6803 字 阅读 →
论文解读

BAT: Better Audio Transformer Guided by Convex Gated Probing

音频分类 | 8.6/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9139 字 阅读 →
论文解读

Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models

语音合成 | 8/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7969 字 阅读 →
论文解读

HyperPotter: Spell the Charm of High-Order Interactions in Audio Deepfake Detection

音频伪造检测 | 7.9/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5063 字 阅读 →
论文解读

NeuroCLUS: A Foundation Model with Functional Clustering for Intracranial Neural Decoding

语音识别 | 6/10

 · 更新于 2026-09-25 · 约 26 分钟 · 12965 字 阅读 →
论文解读

OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention

音视频问答 | 7.5/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7777 字 阅读 →
论文解读

Pianist Transformer: Towards Expressive Piano Performance Rendering via Scalable Self-Supervised Pre-Training

音乐生成 | 8.1/10

 · 更新于 2026-09-25 · 约 6 分钟 · 2789 字 阅读 →
论文解读

Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis

音视频生成 | 7.2/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6404 字 阅读 →
论文解读

SURF: Separation via Unsupervised Remixing Flow

语音分离 | 6.2/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8040 字 阅读 →
论文解读

VocSim A Training-free Benchmark for Zero-shot Content Identity in Single-source Audio

音频检索 | 8.2/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6711 字 阅读 →
论文解读

LMPAN: A Lightweight Multi-Path Alignment Network for Joint Full-Duplex Acoustic Echo Cancellation and Noise Suppression

语音增强 | 6.2/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8743 字 阅读 →
论文解读

Pmeta-TLA: Backdoor Attacks for Speech Classification Models via Meta-Learning with Timbre Leakage Attack

语音唤醒 | 6/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6409 字 阅读 →