论文解读

CORTIS: Text-Only Adaptation of Spoken Language Models for Task-Oriented Voice Agents

语音识别 | 7.7/10

 · 更新于 2026-09-07 · 约 10 分钟 · 4555 字 阅读 →
论文解读

CoughPhase-CLR: Designing an acoustics-informed foundation model for coughing sound classification

对比学习 | 10/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5062 字 阅读 →
论文解读

Cross-lingual Retrieval-Augmented Classification for Dysarthria Severity Assessment

语音识别 | 8.4/10

 · 更新于 2026-09-07 · 约 21 分钟 · 10226 字 阅读 →
论文解读

Direct Raw Audio Signal Processing via Reservoir Computing: An Investigation into 'Feature-Free' Architectures

语音识别 | 4.5/10

 · 更新于 2026-09-07 · 约 9 分钟 · 4329 字 阅读 →
论文解读

DisSpeech: Low-Resource Controllable Mandarin Stuttered Speech Synthesis for ASR Augmentation

语音合成 | 7.2/10

 · 更新于 2026-09-07 · 约 13 分钟 · 6257 字 阅读 →
论文解读

Domain-incremental audio classification using domain-specific experts and prototype classifier

音频分类 | 9/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5508 字 阅读 →
论文解读

Don't Listen to Me: A Lightweight, Low-Latency Model for Own-Voice Cancellation in Far-Field Speech Enhancement

语音增强 | 8.4/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5551 字 阅读 →
论文解读

DSSCNet: A Transfer Learning Framework for Cross-Corpus Dysarthric Speech Severity Classification

迁移学习 | 6.3/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5528 字 阅读 →
论文解读

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis

 · 更新于 2026-09-07 · 约 22 分钟 · 10807 字 阅读 →
论文解读

ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era

语音识别 | 7.5/10

 · 更新于 2026-09-07 · 约 23 分钟 · 11361 字 阅读 →
论文解读

Explainable AI in Speaker Recognition -- Attention Map Visualisation and Evaluation

说话人识别 | 5.5/10

 · 更新于 2026-09-07 · 约 13 分钟 · 6264 字 阅读 →
论文解读

Exploiting Neural Audio Codec Latents for Adversarial Audio Attacks

生成对抗网络 | 7.2/10

 · 更新于 2026-09-07 · 约 10 分钟 · 4695 字 阅读 →
论文解读

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech

 · 更新于 2026-09-07 · 约 23 分钟 · 11512 字 阅读 →
论文解读

From Text Metrics to Model Internals: A Study of Whisper ASR Hallucination Detection

语音识别 | 7.5/10

 · 更新于 2026-09-07 · 约 16 分钟 · 7637 字 阅读 →
论文解读

Gradient-Based Learning of Parametric Engine Sound Representations for Real-Time Resynthesis and Tuning on Embedded Systems

参数高效微调 | 7.8/10

 · 更新于 2026-09-07 · 约 9 分钟 · 4491 字 阅读 →
论文解读

HALAS: A Human-Annotated Dataset of Hallucinations of Modern ASR Systems

语音识别 | 8.4/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5820 字 阅读 →
论文解读

How Well Do Self-Supervised Speech Models Encode Age and Gender in Children's Speech? A Layer-Wise Analysis Across Multiple Architectures

自监督学习 | 9/10

 · 更新于 2026-09-07 · 约 10 分钟 · 4728 字 阅读 →
论文解读

Imitation Learning for Elder-Facing Speech Synthesis

语音合成 | 5.5/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5760 字 阅读 →
论文解读

Improving Engine Sound Analysis in Hot-Test Environments via a RAB-U-Net (Residual Attention Block U-Net) Noise Removal Method

音频降噪 | 4.9/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5465 字 阅读 →
论文解读

Improving Text-to-Music Generation with Human Preference Rewards

音乐生成 | 8.5/10

 · 更新于 2026-09-07 · 约 13 分钟 · 6462 字 阅读 →