论文解读

Benchmarking Commercial ASR Systems on Code-Switching Speech: Arabic, Persian, and German

语音识别 | 6.8/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7727 字 阅读 →
论文解读

Evaluating Speech Articulation Synthesis with Articulatory Phoneme Recognition

语音质量评估 | 8.2/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7861 字 阅读 →
论文解读

FormalASR: End-to-End Spoken Chinese to Formal Text

语音识别 | 8.2/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7790 字 阅读 →
论文解读

Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech

语音合成 | 9.5/10

 · 更新于 2026-09-25 · 约 20 分钟 · 9678 字 阅读 →
论文解读

SEABAD: A Tropical Bird Activity Detection Dataset for Passive Acoustic Monitoring

生物声学 音频事件检测 | 8.1/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8250 字 阅读 →
论文解读

FormalASR: End-to-End Spoken Chinese to Formal Text

语音识别 | 6/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7432 字 阅读 →
论文解读

GroupAffect-4: A Multimodal Dataset of Four-Person Collaborative Interaction

数据集 | 6.8/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9307 字 阅读 →
论文解读

Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation

语音识别 | 6.8/10

 · 更新于 2026-09-25 · 约 22 分钟 · 10575 字 阅读 →
论文解读

Audio-Image Cross-Modal Retrieval with Onomatopoeic Images

音频检索 | 7/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6984 字 阅读 →
论文解读

Beyond Transcripts: Iterative Peer-Editing with Audio Unlocks High-Quality Human Summaries of Conversational Speech

语音摘要 | 7.2/10

 · 更新于 2026-09-25 · 约 23 分钟 · 11227 字 阅读 →
论文解读

Profiling the Voice: Speaker-Specific Phoneme Fingerprinting for Speech Deepfake Detection

语音伪造检测 | 7/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7180 字 阅读 →
论文解读

Sonalyzer-Moz: A Framework for Analyzing the Structure of Mozart's Sonata Form

音乐结构分析 | 7.3/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6193 字 阅读 →
论文解读

UrduSpeech: A 156-Hour Urdu Speech Corpus with 12-Dimension Paralinguistic Annotations

语音识别 | 7/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7835 字 阅读 →
论文解读

WavFlow: Audio Generation in Waveform Space

音频生成 | 6.7/10

 · 更新于 2026-09-25 · 约 24 分钟 · 11712 字 阅读 →
论文解读

FSD50K-Solo: Automated Curation of Single-Source Sound Events

数据清洗 | 5.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6839 字 阅读 →
论文解读

IsoNet: Spatially-aware audio-visual target speech extraction in complex acoustic environments

语音提取 | 6/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8685 字 阅读 →
论文解读

Masked Autoencoders with Limited Data: Does It Work? A Fine-Grained Bioacoustics Case Study

音频分类 | 5.5/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9405 字 阅读 →
论文解读

PROCESS-2: A Benchmark Speech Corpus for Early Cognitive Impairment Detection

语音生物标志物 | 5.4/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8039 字 阅读 →
论文解读

Towards Fine-Grained Multi-Dimensional Speech Understanding: Data Pipeline, Benchmark, and Model

Towards Fine-Grained Multi-Dimensional Speech Understanding: Data Pipeline, Benchmark, and Model

 · 更新于 2026-09-25 · 约 18 分钟 · 8956 字 阅读 →
论文解读

What makes a word hard to learn? Modeling L1 influence on English vocabulary difficulty

词汇难度预测 | 5.0/10

 · 更新于 2026-09-25 · 约 20 分钟 · 9671 字 阅读 →