论文解读

EchoHawk: A Reproducible Acoustic Pipeline for Drone Detection, Classification, and Direction-Finding, with a Cautionary Study of Session-Level Data Leakage

EchoHawk: A Reproducible Acoustic Pipeline for Drone Detection, Classification, and Direction-Finding, with a Cautionary Study of Session-Level Data Leakage

 · 更新于 2026-09-06 · 约 9 分钟 · 4133 字 阅读 →
论文解读

Effective Depth in Joint Source-Channel Coding: An Implicit Equilibrium Analysis

语音编码 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4439 字 阅读 →
论文解读

Evaluation of Head-Related Transfer Functions Across Five Levels of Individualisation in Virtual Reality

空间音频 | 7.6/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5571 字 阅读 →
论文解读

FacePlex: Full-Duplex Joint Speech-Facial Motion Generation for Conversational Avatars

语音合成 | 7.8/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6765 字 阅读 →
论文解读

GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark

语音识别 | 8.7/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5725 字 阅读 →
论文解读

How to Leverage Synthetic Speech for LLM-Based ASR Systems?

语音识别 | 8.7/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7623 字 阅读 →
论文解读

Improving Large-Scale Weakly Supervised ASR by Filtering and Selection

Improving Large-Scale Weakly Supervised ASR by Filtering and Selection

 · 更新于 2026-09-06 · 约 22 分钟 · 10708 字 阅读 →
论文解读

LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training

音乐生成 | 9.4/10

 · 更新于 2026-09-06 · 约 4 分钟 · 1658 字 阅读 →
论文解读

LoRA-Tuned Large Language Models for Dementia Detection via Multi-View Speech-Derived Features

参数高效微调 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4940 字 阅读 →
论文解读

MeloDISinger: Melody-Aware & Duration-Preserving Singing Voice Editing with Audio Infilling

语音合成 | 7.4/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6617 字 阅读 →
论文解读

OLIVE: View-Augmented Latent Prediction with Waveform Reconstruction for Speech SSL

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7900 字 阅读 →
论文解读

Position-Aware Target Speaker Extraction for Long-Form Multi-Party Conversations: A Diarization-Free Framework for ASR

语音识别 | 8.5/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6087 字 阅读 →
论文解读

Predicting Timbre Traits for Interpretable Assessment of Musical Sound Synthesizers

音频生成 | 6.1/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5084 字 阅读 →
论文解读

Preference-ASR: A Preference-Aware Test Set for Benchmarking ASR in the Era of Speech LLMs

语音识别 | 9.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5833 字 阅读 →
论文解读

Proteus: Automated Adversarial Robustness Testing for Audio Deepfake Detectors

数据增强 | 5.3/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4768 字 阅读 →
论文解读

Rehearsed Multi-Agent Live Product Demonstrations with Real-Time Voice Question Answering

多模态模型 | 5.3/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6119 字 阅读 →
论文解读

Semi-Supervised Sound Event Detection with Conditional Mixup and Embedding-Level Contrastive Loss

对比学习 | 7.6/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5962 字 阅读 →
论文解读

SICAGE: Speaker-Independent Culture-Aware Gesture Generation using TED4C-L Dataset

语音合成 | 8.7/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6507 字 阅读 →
论文解读

SIGMA: Saliency-Guided Sparse Mask Attacks for Speech Emotion Recognition

语音情感识别 | 7.1/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5240 字 阅读 →
论文解读

SIMAX: A Scalable and Interpretable Framework for Multi-Fidelity and Annotated Clinician-Patient Dialogue Simulation

语音合成 | 6.6/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6319 字 阅读 →