论文解读

Modeling Both Intra- And Inter-Utterance Variability for Conversational Emotion Recognition

语音情感识别 | 6.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4447 字 阅读 →
论文解读

Multimodal LLMs as Expert Speech Annotators: Acoustic Macro-Descriptors for Parkinson's Detection

语音生物标志物 | 6.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4079 字 阅读 →
论文解读

PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models

语音对话系统 | 8.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4703 字 阅读 →
论文解读

PFluxTTS: Hybrid Flow-Matching TTS with Robust Cross-Lingual Voice Cloning and Inference-Time Model Fusion

语音合成 | 7.0/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6124 字 阅读 →
论文解读

Plug-and-Play Emotion Graphs for Compositional Prompting in Zero-Shot Speech Emotion Recognition

语音情感识别 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4538 字 阅读 →
论文解读

Poly-SVC: Polyphony-Aware Singing Voice Conversion with Harmonic Modeling

歌唱语音转换 | 6.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4949 字 阅读 →
论文解读

Probing the Hidden Talent of ASR foundation models for L2 English Oral Assessment

预训练 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4262 字 阅读 →
论文解读

QE-XVC: Zero-Shot Cross-Lingual Voice Conversion via Query-Enhancement and Conditional Flow Matching

语音转换 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4561 字 阅读 →
论文解读

RFM-Editing: Rectified Flow Matching for Text-Guided Audio Editing

音频编辑 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4151 字 阅读 →
论文解读

Salad-VAE: Semantic Audio Compression with Language-Audio Distillation

音频压缩 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4645 字 阅读 →
论文解读

Separate this, and all of these Things Around It: Music Source Separation Via Hyperellipsoidal Queries

音乐分离 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4838 字 阅读 →
论文解读

Sing What You Fit: A Perception-Based Dataset and Benchmark for Vocal-Song Suitability Analysis

音乐信息检索 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 5002 字 阅读 →
论文解读

SmoothCLAP: Soft-Target Enhanced Contrastive Language-Audio Pretraining for Affective Computing

语音情感识别 | 6.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3757 字 阅读 →
论文解读

SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4626 字 阅读 →
论文解读

SpeechMapper: Speech-To-Text Embedding Projector for LLMs

语音大模型 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5397 字 阅读 →
论文解读

Spiking Attention Network: A Hybrid Neuromorphic Approach to Underwater Acoustic Localization and Zero-Shot Adaptation

声源定位 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3981 字 阅读 →
论文解读

Spiking Temporal-Enhanced Network for Zero-Shot Audio-Visual Learning

音频分类 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4341 字 阅读 →
论文解读

StylePitcher: Generating Style-Following and Expressive Pitch Curves for Versatile Singing Tasks

歌唱语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5251 字 阅读 →
论文解读

Synthesized Data Selection via Score Distribution Matching for Te Reo Māori Automatic Speech Recognition

语音识别 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4219 字 阅读 →
论文解读

T-Cache: Fast Inference For Masked Generative Transformer-Based TTS Via Prompt-Aware Feature Caching

语音合成 | 9.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4509 字 阅读 →