论文解读

AdaTT: Text-Guided Instrument Timbre Transfer with Target-Adaptive Structural Control

音频生成 | 8.7/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5380 字 阅读 →
论文解读

Time-frequency localization of bird calls in dense soundscapes

信号处理基础 | 8.5/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6377 字 阅读 →
论文解读

Acoustic Cue Alignment in Audio Language Models for Speech Emotion Recognition

语音情感识别 | 7.8/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5285 字 阅读 →
论文解读

How Far Can Chord-Symbol Time-Series Adaptation Carry Genre Identity? Capabilities and Boundaries in Multi-Genre Chord-Symbol Modeling

音乐信息检索 | 8.8/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5729 字 阅读 →
论文解读

MyGardenBird: A Machine-Learning-Ready Bird Sound Dataset for Twelve Common Malaysian Birds

音频事件检测 | 7.2/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5600 字 阅读 →
论文解读

Phonetic Error Analysis of Raw Waveform Acoustic Models

语音识别 | 7.6/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4614 字 阅读 →
论文解读

Multilingual Detection of Alzheimer's Disease from Speech: A Cross-Linguistic Transfer Learning Approach

迁移学习 | 5.7/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4489 字 阅读 →
论文解读

Sound Effects Dataset Unification With the Universal Category System

音频分类 | 6.9/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5518 字 阅读 →
论文解读

USAD 2.0: Scaling Representation Distillation for Universal Audio Understanding

音频编码 | 9/10

 · 更新于 2026-09-24 · 约 12 分钟 · 6009 字 阅读 →
论文解读

BaltiVoice: A Speech Corpus and Fine-tuned Whisper ASR System for the Balti Language

语音识别 | 7.8/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4606 字 阅读 →
论文解读

Improving acoustic drone detection generalization through pretraining and data augmentation

音频事件检测 | 7.7/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5246 字 阅读 →
论文解读

A Multi-Probe Audit of Clinical-Interview Depression Detection Benchmarks

语音情感识别 | 9.6/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6505 字 阅读 →
论文解读

Cost-Effective Model Evaluation with Meta-Learning

迁移学习 | 5.4/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7948 字 阅读 →
论文解读

Diffusion Domain Expansion: Learning to Coordinate Pre-trained Diffusion Models

扩散模型 | 7.4/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5697 字 阅读 →
论文解读

Natural Yet Challenging to Detect: Robust In-the-Wild TTS through EMA and Dual-Scoring Prompt Selection -- Submission for WildSpoof 2026 TTS Track

语音合成 | 5.2/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5445 字 阅读 →
论文解读

CoarseSoundNet: Building a reliable model for ecological soundscape analysis

音频分类 | 8.5/10

 · 更新于 2026-09-24 · 约 17 分钟 · 8312 字 阅读 →
论文解读

Audio-Image Cross-Modal Retrieval with Onomatopoeic Images

音频检索 | 7/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6984 字 阅读 →
论文解读

Break-the-Beat! Controllable MIDI-to-Drum Audio Synthesis

音频生成 | 6.8/10

 · 更新于 2026-09-24 · 约 17 分钟 · 8349 字 阅读 →
论文解读

Persian MusicGen: A Large-Scale Dataset and Culturally-Aware Generative Model for Persian Music

音乐生成 | 6.7/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6651 字 阅读 →
论文解读

Refining Pseudo-Audio Prompts with Speech-Text Alignment for Text-Only Domain Adaptation in LLM-Based ASR

语音识别 | 7.5/10

 · 更新于 2026-09-24 · 约 17 分钟 · 8162 字 阅读 →