论文解读

A Finetuned SpeechLLM for Joint Multi-Granular L2 Assessment and Natural-Language Rationales

大语言模型 | 10/10

 · 更新于 2026-09-09 · 约 16 分钟 · 7516 字 阅读 →
论文解读

A Hierarchical Feature Engineering Framework for Automated Classification of Phonotraumatic and Non-Phonotraumatic Vocal Hyperfunction

A Hierarchical Feature Engineering Framework for Automated Classification of Phonotraumatic and Non-Phonotraumatic Vocal Hyperfunction

 · 更新于 2026-09-09 · 约 12 分钟 · 5753 字 阅读 →
论文解读

A study on the impact of region specific data on the performance of Indic ASR

语音识别 | 7.2/10

 · 更新于 2026-09-09 · 约 13 分钟 · 6491 字 阅读 →
论文解读

AeroSpectra Sentinel: An Auditable LLM Prompt-Chaining Decision-Support Workflow for Acute Asthma Risk Assessment from Respiratory Sounds and Clinical Signals

音频事件检测 | 4.5/10

 · 更新于 2026-09-09 · 约 11 分钟 · 5227 字 阅读 →
论文解读

Assessing the Energy and Carbon Emissions of Neural Speaker Verification Model in Training and Inference

说话人验证 | 7.4/10

 · 更新于 2026-09-09 · 约 12 分钟 · 5733 字 阅读 →
论文解读

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs

语音识别 | 8.8/10

 · 更新于 2026-09-09 · 约 13 分钟 · 6046 字 阅读 →
论文解读

BareWave: Waveform-Native Flow-Matching Text-to-Speech

语音合成 | 7.0/10

 · 更新于 2026-09-09 · 约 17 分钟 · 8292 字 阅读 →
论文解读

Bridging Traditional Explainability Methods and Multimodal Multilingual Models: An XAI-Based Analysis

语音识别 | 5.4/10

 · 更新于 2026-09-09 · 约 18 分钟 · 8872 字 阅读 →
论文解读

Can LLMs understand LilyPond? A benchmark for symbolic music generation and understanding

音乐生成 | 7/10

 · 更新于 2026-09-09 · 约 11 分钟 · 5101 字 阅读 →
论文解读

Conan-embedding-v3: Fusing Modality-Specific Models for Omni-Modal Embedding

音频检索 | 7.6/10

 · 更新于 2026-09-09 · 约 12 分钟 · 5579 字 阅读 →
论文解读

Cross-Modal Masking for Robust Silent Speech Synthesis Using sEMG and Lipreading

语音合成 | 7.5/10

 · 更新于 2026-09-09 · 约 16 分钟 · 7716 字 阅读 →
论文解读

Discovering Functionally Selective Brain Regions with a Deep Topographic Multimodal Model

多模态模型 | 8.2/10

 · 更新于 2026-09-09 · 约 11 分钟 · 5464 字 阅读 →
论文解读

End-to-End Training for Discrete Token LLM based TTS System

语音合成 | 7.6/10

 · 更新于 2026-09-09 · 约 14 分钟 · 6973 字 阅读 →
论文解读

Exploring the Scale and Diversity of Speech Anti-spoofing Datasets: Experiments and Analysis

数据增强 | 7.4/10

 · 更新于 2026-09-09 · 约 17 分钟 · 8394 字 阅读 →
论文解读

Factors affecting ASR performance: A study using state of the art ASR models in Indic Languages

语音识别 | 6.9/10

 · 更新于 2026-09-09 · 约 12 分钟 · 5995 字 阅读 →
论文解读

Fast and Robust On-Device Speaker Diarization: Relative Minimum Cluster Size for Stride-Accelerated Pipelines

说话人分离 | 6.6/10

 · 更新于 2026-09-09 · 约 12 分钟 · 5839 字 阅读 →
论文解读

Few-shot Class-variable Incremental Audio Classification via Prototype Adaptation and Pseudo Class-variable Training

音频分类 | 6.9/10

 · 更新于 2026-09-09 · 约 12 分钟 · 5750 字 阅读 →
论文解读

FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation

语音合成 | 7.9/10

 · 更新于 2026-09-09 · 约 10 分钟 · 4750 字 阅读 →
论文解读

From A to B to A: Palindromic Zero-Shot Voice Conversion with Non-Parallel Data

From A to B to A: Palindromic Zero-Shot Voice Conversion with Non-Parallel Data

 · 更新于 2026-09-09 · 约 25 分钟 · 12306 字 阅读 →
论文解读

FXplorer: A Map-Based Interface for Exploratory Audio Effect Design

音频生成 | 7.5/10

 · 更新于 2026-09-09 · 约 10 分钟 · 4968 字 阅读 →