论文解读

Multi-Task Learning For Speech Quality Assessment Using ASR-Derived Entropy Features

语音质量评估 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4836 字 阅读 →
论文解读

Multilingual Supervised Pretraining with Lm-Assisted Decoding for Visual Speech Recognition

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3573 字 阅读 →
论文解读

Multimodal Transformer with Multiperspective Training for Predicting Self-Expression Skills from Video Interview

多模态模型 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4288 字 阅读 →
论文解读

MuseTok: Symbolic Music Tokenization for Generation and Semantic Understanding

音乐生成 | 8.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5119 字 阅读 →
论文解读

Noise-Robust AV-ASR Using Visual Features both in the Whisper Encoder and Decoder

语音识别 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4647 字 阅读 →
论文解读

On deepfake voice detection - It’s all in the presentation

音频深度伪造检测 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4174 字 阅读 →
论文解读

Online Register For Dual-Mode Self-Supervised Speech Models: Mitigating the Lack of Future Context

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5121 字 阅读 →
论文解读

PADAM: Perceptual Audio Defect Assessment Model

音频分类 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4814 字 阅读 →
论文解读

Probing the Hidden Talent of ASR foundation models for L2 English Oral Assessment

预训练 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4262 字 阅读 →
论文解读

Quality Assessment of Noisy and Enhanced Speech with Limited Data: UWB-NTIS System for Voicemos 2024

语音质量评估 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5174 字 阅读 →
论文解读

RASD-SR: A Robust Anomalous Sound Detection Framework with Score Recalibration

异常声音检测 | 8.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5155 字 阅读 →
论文解读

Reading Between the Waves: Robust Topic Segmentation Using Inter-Sentence Audio Features

音频分类 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5113 字 阅读 →
论文解读

Reasoning Driven Captions to Assist Noise Robust Speech Emotion Recognition

语音情感识别 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4690 字 阅读 →
论文解读

Recovering Performance in Speech Emotion Recognition from Discrete Tokens Via Multi-Layer Fusion and Paralinguistic Feature Integration

语音情感识别 | 6.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4949 字 阅读 →
论文解读

Reference-Aware SFM Layers for Intrusive Intelligibility Prediction

语音评估 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4990 字 阅读 →
论文解读

SAASDNet: An EEG-Based Streaming Auditory Attention Switch Decoding Network for Self-Initiated Attention Switching in Mixed Speech

脑机接口 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4824 字 阅读 →
论文解读

SAUNA: Song-Level Audio & User-Listening Data Neural Alignment

音乐信息检索 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3885 字 阅读 →
论文解读

Scaling Multi-Talker ASR with Speaker-Agnostic Activity Streams

语音识别 | 8.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4626 字 阅读 →
论文解读

SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper

语音识别 | 8.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5790 字 阅读 →
论文解读

Shared Representation Learning for Reference-Guided Targeted Sound Detection

音频事件检测 | 8.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4764 字 阅读 →