论文解读

AppTek Call-Center Dialogues: A Multi-Accent Long-Form Benchmark for English ASR

语音识别 | 6.5/10

 · 更新于 2026-10-02 · 约 9 分钟 · 4392 字 阅读 →
论文解读

Audio Effect Estimation with DNN-Based Prediction and Search Algorithm

音乐理解 | 8.0/10

 · 更新于 2026-10-02 · 约 10 分钟 · 4814 字 阅读 →
论文解读

Audio Video Verbal Analysis (AVVA) for Capturing Classroom Dialogues

音频问答 | 6.0/10

 · 更新于 2026-10-02 · 约 9 分钟 · 4011 字 阅读 →
论文解读

Beyond Acoustic Sparsity and Linguistic Bias: A Prompt-Free Paradigm for Mispronunciation Detection and Diagnosis

发音错误检测 | 8.5/10

 · 更新于 2026-10-02 · 约 13 分钟 · 6451 字 阅读 →
论文解读

Beyond the Baseband: Adaptive Multi-Band Encoding for Full-Spectrum Bioacoustics Classification

生物声学 | 7.0/10

 · 更新于 2026-10-02 · 约 13 分钟 · 6321 字 阅读 →
论文解读

BUT System Description for CHiME-9 MCoRec Challenge

语音识别 | 6.5/10

 · 更新于 2026-10-02 · 约 11 分钟 · 5308 字 阅读 →
论文解读

DM-ASR: Diarization-aware Multi-speaker ASR with Large Language Models

说话人识别 | 8.0/10

 · 更新于 2026-10-02 · 约 10 分钟 · 4581 字 阅读 →
论文解读

Do Sparse Autoencoders Capture Concept Manifolds?

可解释性 | 7.0/10

 · 更新于 2026-10-02 · 约 15 分钟 · 7300 字 阅读 →
论文解读

Dual-LoRA: Parameter-Efficient Adversarial Disentanglement for Cross-Lingual Speaker Verification

说话人验证 | 7.0/10

 · 更新于 2026-10-02 · 约 11 分钟 · 5273 字 阅读 →
论文解读

Earable Platform with Integrated Simultaneous EEG Sensing and Auditory Stimulation

音频事件检测 | 5.5/10

 · 更新于 2026-10-02 · 约 6 分钟 · 2641 字 阅读 →
论文解读

EdgeSpike: Spiking Neural Networks for Low-Power Autonomous Sensing in Edge IoT Architectures

音频事件检测 | 7.5/10

 · 更新于 2026-10-02 · 约 16 分钟 · 7624 字 阅读 →
论文解读

Few-Shot Accent Synthesis for ASR with LLM-Guided Phoneme Editing

语音识别 | 7.5/10

 · 更新于 2026-10-02 · 约 11 分钟 · 5168 字 阅读 →
论文解读

Full-Duplex Interaction in Spoken Dialogue Systems: A Comprehensive Study from the ICASSP 2026 HumDial Challenge

语音对话系统 | 6.5/10

 · 更新于 2026-10-02 · 约 8 分钟 · 3578 字 阅读 →
论文解读

HATS: An Open data set Integrating Human Perception Applied to the Evaluation of Automatic Speech Recognition Metrics

语音识别 | 7.0/10

 · 更新于 2026-10-02 · 约 8 分钟 · 3963 字 阅读 →
论文解读

Identifying and typifying demographic unfairness in phoneme-level embeddings of self-supervised speech recognition models

语音识别 | 7.0/10

 · 更新于 2026-10-02 · 约 10 分钟 · 4560 字 阅读 →
论文解读

JaiTTS: A Thai Voice Cloning Model

语音合成 | 7.5/10

 · 更新于 2026-10-02 · 约 5 分钟 · 2500 字 阅读 →
论文解读

Listening with Time: Precise Temporal Awareness for Long-Form Audio Understanding

音频场景理解 | 8.0/10

 · 更新于 2026-10-02 · 约 10 分钟 · 4835 字 阅读 →
论文解读

LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition

语音识别 | 9.0/10

 · 更新于 2026-10-02 · 约 9 分钟 · 4128 字 阅读 →
论文解读

Mapping the Methodological Space of Classroom Interaction Research: Scale, Duration, and Modality in an Age of AI

模型评估 | 6.0/10

 · 更新于 2026-10-02 · 约 6 分钟 · 2848 字 阅读 →
论文解读

MCPHunt: An Evaluation Framework for Cross-Boundary Data Propagation in Multi-Server MCP Agents

模型评估 | 7.5/10

 · 更新于 2026-10-02 · 约 13 分钟 · 6271 字 阅读 →