论文解读

NCF-TTS: Enhancing Flow Matching Based Text-To-Speech with Neighborhood Consistency Flow

语音合成 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4168 字 阅读 →
论文解读

Neural Network-Based Time-Frequency-Bin-Wise Linear Combination of Beamformers for Underdetermined Target Source Extraction

语音分离 | 7.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5398 字 阅读 →
论文解读

Neuromamba: Adaptive Frequency Filtering with a Pyramid Mamba for sEEG-driven Speech Synthesis

语音合成 | 8.0/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5962 字 阅读 →
论文解读

NeuroSIFT: A Biologically-Inspired Framework with Explicit Signal-Noise Separation for Robust Multimodal Emotion Recognition

多模态情感识别 | 8.0/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3891 字 阅读 →
论文解读

nGPT as a Scalable Architecture for Speech Recognition and Translation

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5927 字 阅读 →
论文解读

No Verifiable Reward for Prosody: Toward Preference-Guided Prosody Learning in TTS

语音合成 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4301 字 阅读 →
论文解读

Noise-Robust AV-ASR Using Visual Features both in the Whisper Encoder and Decoder

语音识别 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4647 字 阅读 →
论文解读

Noise-Robust Contrastive Learning with an MFCC-Conformer for Coronary Artery Disease Detection

音频分类 | 7.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5224 字 阅读 →
论文解读

Noise-to-Notes: Diffusion-Based Generation and Refinement for Automatic Drum Transcription

音乐信息检索 | 8.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5198 字 阅读 →
论文解读

Non-Line-of-Sight Vehicle Detection via Audio-Visual Fusion

音频分类 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4086 字 阅读 →
论文解读

Obstructive Sleep Apnea Endotype Prediction During Wakefulness Using Voice Biomarkers

语音生物标志物 | 6.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3906 字 阅读 →
论文解读

Off-The-Grid Multi-Pitch Estimation Using Optimal Transport

音乐信息检索 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4885 字 阅读 →
论文解读

OMNI-AVSR: Towards Unified Multimodal Speech Recognition With Large Language Models

语音识别 | 8.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4649 字 阅读 →
论文解读

On deepfake voice detection - It’s all in the presentation

音频深度伪造检测 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4174 字 阅读 →
论文解读

On The Design of Efficient Neural Methods for Geometry-Agnostic Multichannel Speech Enhancement

语音增强 | 6.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4364 字 阅读 →
论文解读

On the Design of Higher-Order Time-Intensity Microphone Arrays for Panoramic Audio Recording and Reproduction

空间音频 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4217 字 阅读 →
论文解读

One Model–Three Tasks: Discovering a Shared Winning Ticket for Low-Complexity Audio Intelligence

音频分类 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3730 字 阅读 →
论文解读

Online Register For Dual-Mode Self-Supervised Speech Models: Mitigating the Lack of Future Context

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5121 字 阅读 →
论文解读

Optimizing Domain-Adaptive Self-Supervised Learning for Clinical Voice-Based Disease Classification

语音生物标志物 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4984 字 阅读 →
论文解读

Optimizing Speech Language Models for Acoustic Consistency

语音合成 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4747 字 阅读 →