论文解读

Efficient Punctuation Restoration via Weighted Lookahead Scoring Method for Streaming ASR Systems

大语言模型 | 6.3/10

 · 更新于 2026-09-09 · 约 10 分钟 · 4987 字 阅读 →
论文解读

Enhancing Audio Captioning with Auxiliary AudioSet Semantics

Enhancing Audio Captioning with Auxiliary AudioSet Semantics

 · 更新于 2026-09-09 · 约 11 分钟 · 5377 字 阅读 →
论文解读

Exploring LLMs for South Asian Music Understanding and Generation

音乐生成 | 7.7/10

 · 更新于 2026-09-09 · 约 11 分钟 · 5025 字 阅读 →
论文解读

F3-Tokenizer: Taming Audio Autoencoder Latents for Understanding and Generation

语音合成 | 7.2/10

 · 更新于 2026-09-09 · 约 13 分钟 · 6113 字 阅读 →
论文解读

FiLM-Based Speaker Conditioning of a SpeechLLM for Pathological Speech Recognition

语音识别 | 8.0/10

 · 更新于 2026-09-09 · 约 15 分钟 · 7310 字 阅读 →
论文解读

FoeGlass: Simple In-Context Learning Is Enough for Red Teaming Audio Deepfake Detectors

音频生成 | 7.5/10

 · 更新于 2026-09-09 · 约 17 分钟 · 8421 字 阅读 →
论文解读

Forgive or forget: Understanding the context of hate in audio retrieval systems

音频检索 | 7.4/10

 · 更新于 2026-09-09 · 约 13 分钟 · 6360 字 阅读 →
论文解读

FORTE: FOL-guided Optimal Refinement for Text-audio rEtrieval

参数高效微调 | 8.1/10

 · 更新于 2026-09-09 · 约 14 分钟 · 6667 字 阅读 →
论文解读

GLASS: GRPO-Trained LoRA for Acoustic Style Steering in Zero-Shot Text-to-Speech

语音合成 | 8.2/10

 · 更新于 2026-09-09 · 约 14 分钟 · 6599 字 阅读 →
论文解读

InfoShield: Privacy-Preserving Speech Representations for Mental Health Screening via Information-Theoretic Optimization

InfoShield: Privacy-Preserving Speech Representations for Mental Health Screening via Information-Theoretic Optimization

 · 更新于 2026-09-09 · 约 15 分钟 · 7171 字 阅读 →
论文解读

Learning Emotion-discriminative Representations for Zero-Shot Cross-lingual Speech Emotion Recognition

语音情感识别 | 8.1/10

 · 更新于 2026-09-09 · 约 12 分钟 · 5728 字 阅读 →
论文解读

M2S-AVSR: Modality-aware Multi-view Self-supervised Representation for Robust Audio-Visual Speech Recognition

语音识别 | 9/10

 · 更新于 2026-09-09 · 约 13 分钟 · 6425 字 阅读 →
论文解读

MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models

语音识别 | 8.4/10

 · 更新于 2026-09-09 · 约 14 分钟 · 6696 字 阅读 →
论文解读

Multi-task Learning is Not Enough: Representational Entanglement in Dual-output Second Language Speech Recognition

语音识别 | 6.9/10

 · 更新于 2026-09-09 · 约 10 分钟 · 4724 字 阅读 →
论文解读

Multilingual Detection of Alzheimer's Disease from Speech: A Cross-Linguistic Transfer Learning Approach

迁移学习 | 5.7/10

 · 更新于 2026-09-09 · 约 9 分钟 · 4489 字 阅读 →
论文解读

nnAudio 2: Overcoming Dynamic Compilation Barriers and Transform Inconsistencies

开源工具 | 7.5/10

 · 更新于 2026-09-09 · 约 12 分钟 · 5846 字 阅读 →
论文解读

Ouvia: A User-centered Framework for Measuring Usability of Speech Translation in Real-World Communication Scenarios

语音翻译 | 8.6/10

 · 更新于 2026-09-09 · 约 12 分钟 · 5825 字 阅读 →
论文解读

Probing Spatial Structure in Pretrained Audio Representations

Probing Spatial Structure in Pretrained Audio Representations

 · 更新于 2026-09-09 · 约 9 分钟 · 4351 字 阅读 →
论文解读

ProSarc: Prosody-Aware Sarcasm Recognition Framework via Temporal Prosodic Incongruity

语音情感识别 | 7.5/10

 · 更新于 2026-09-09 · 约 14 分钟 · 6641 字 阅读 →
论文解读

Revisiting Lexicon Evaluation in Unsupervised Word Discovery

语音识别 | 1/10

 · 更新于 2026-09-09 · 约 11 分钟 · 5346 字 阅读 →