论文解读

Do Sparse Autoencoders Capture Concept Manifolds?

可解释性 | 7.0/10

 · 更新于 2026-09-10 · 约 15 分钟 · 7300 字 阅读 →
论文解读

Dual-LoRA: Parameter-Efficient Adversarial Disentanglement for Cross-Lingual Speaker Verification

说话人验证 | 7.0/10

 · 更新于 2026-09-10 · 约 11 分钟 · 5273 字 阅读 →
论文解读

Earable Platform with Integrated Simultaneous EEG Sensing and Auditory Stimulation

音频事件检测 | 5.5/10

 · 更新于 2026-09-10 · 约 6 分钟 · 2641 字 阅读 →
论文解读

EdgeSpike: Spiking Neural Networks for Low-Power Autonomous Sensing in Edge IoT Architectures

音频事件检测 | 7.5/10

 · 更新于 2026-09-10 · 约 16 分钟 · 7624 字 阅读 →
论文解读

Few-Shot Accent Synthesis for ASR with LLM-Guided Phoneme Editing

语音识别 | 7.5/10

 · 更新于 2026-09-10 · 约 11 分钟 · 5168 字 阅读 →
论文解读

Full-Duplex Interaction in Spoken Dialogue Systems: A Comprehensive Study from the ICASSP 2026 HumDial Challenge

语音对话系统 | 6.5/10

 · 更新于 2026-09-10 · 约 8 分钟 · 3578 字 阅读 →
论文解读

HATS: An Open data set Integrating Human Perception Applied to the Evaluation of Automatic Speech Recognition Metrics

语音识别 | 7.0/10

 · 更新于 2026-09-10 · 约 8 分钟 · 3963 字 阅读 →
论文解读

Identifying and typifying demographic unfairness in phoneme-level embeddings of self-supervised speech recognition models

语音识别 | 7.0/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4560 字 阅读 →
论文解读

JaiTTS: A Thai Voice Cloning Model

语音合成 | 7.5/10

 · 更新于 2026-09-10 · 约 5 分钟 · 2500 字 阅读 →
论文解读

Listening with Time: Precise Temporal Awareness for Long-Form Audio Understanding

音频场景理解 | 8.0/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4835 字 阅读 →
论文解读

LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition

语音识别 | 9.0/10

 · 更新于 2026-09-10 · 约 9 分钟 · 4128 字 阅读 →
论文解读

Mapping the Methodological Space of Classroom Interaction Research: Scale, Duration, and Modality in an Age of AI

模型评估 | 6.0/10

 · 更新于 2026-09-10 · 约 6 分钟 · 2848 字 阅读 →
论文解读

MCPHunt: An Evaluation Framework for Cross-Boundary Data Propagation in Multi-Server MCP Agents

模型评估 | 7.5/10

 · 更新于 2026-09-10 · 约 13 分钟 · 6271 字 阅读 →
论文解读

MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction

语音对话系统 | 8.5/10

 · 更新于 2026-09-10 · 约 14 分钟 · 6664 字 阅读 →
论文解读

Normativity and Productivism: Ableist Intelligence? A Degrowth Analysis of AI Sign Language Translation Tools for Deaf People

语音翻译 | 3.5/10

 · 更新于 2026-09-10 · 约 5 分钟 · 2038 字 阅读 →
论文解读

Predicting Upcoming Stuttering Events from Three-Second Audio: Stratified Evaluation Reveals Severity-Selective Precursors, and the Model Deploys Fully On-Device

语音生物标志物 | 7.0/10

 · 更新于 2026-09-10 · 约 13 分钟 · 6353 字 阅读 →
论文解读

Qualitative Evaluation of Language Model Rescoring in Automatic Speech Recognition

语音识别 | 6.5/10

 · 更新于 2026-09-10 · 约 5 分钟 · 2252 字 阅读 →
论文解读

Selective Augmentation: Improving Universal Automatic Phonetic Transcription via G2P Bootstrapping

语音识别 | 6.0/10

 · 更新于 2026-09-10 · 约 7 分钟 · 3160 字 阅读 →
论文解读

Spectrographic Portamento Gradient Analysis: A Quantitative Method for Historical Cello Recordings with Application to Beethoven's Piano and Cello Sonatas, 1930--2012

音乐信息检索 | 7.5/10

 · 更新于 2026-09-10 · 约 7 分钟 · 3274 字 阅读 →
论文解读

Taming Noise-Induced Prototype Degradation for Privacy-Preserving Personalized Federated Fine-Tuning

个性化联邦学习 | 7.5/10

 · 更新于 2026-09-10 · 约 7 分钟 · 3077 字 阅读 →