论文解读

VocSim A Training-free Benchmark for Zero-shot Content Identity in Single-source Audio

音频检索 | 8.2/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6711 字 阅读 →
论文解读

CosyEdit2: Speech-Editing-Oriented Reinforcement Learning Unlocks Better Zero-Shot TTS

语音合成 | 8/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6277 字 阅读 →
论文解读

CodeBind: Decoupled Representation Learning for Multimodal Alignment with Unified Compositional Codebook

多模态模型 | 8.6/10

 · 更新于 2026-09-24 · 约 21 分钟 · 10221 字 阅读 →
论文解读

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis

语音合成 | 6.8/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7239 字 阅读 →
论文解读

NAACA: Training-Free NeuroAuditory Attentive Cognitive Architecture with Oscillatory Working Memory for Salience-Driven Attention Gating

音频事件检测 | 7.0/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7828 字 阅读 →
论文解读

Exploring Token-Space Manipulation in Latent Audio Tokenizers

音频编码 | 6.5/10

 · 更新于 2026-09-24 · 约 19 分钟 · 9238 字 阅读 →
论文解读

Poly-SVC: Polyphony-Aware Singing Voice Conversion with Harmonic Modeling

歌唱语音转换 | 5.5/10

 · 更新于 2026-09-24 · 约 28 分钟 · 13767 字 阅读 →
论文解读

Polyphonia: Zero-Shot Timbre Transfer in Polyphonic Music with Acoustic-Informed Attention Calibration

音乐生成 | 7.5/10

 · 更新于 2026-09-24 · 约 18 分钟 · 8830 字 阅读 →
论文解读

Zero-Shot Imagined Speech Decoding via Imagined-to-Listened MEG Mapping

Zero-Shot Imagined Speech Decoding via Imagined-to-Listened MEG Mapping

 · 更新于 2026-09-24 · 约 13 分钟 · 6248 字 阅读 →
论文解读

Automated Clinical Report Generation for Remote Cognitive Remediation: Comparing Knowledge-Engineered Templates and LLMs in Low-Resource Settings

临床报告生成 | 7.5/10

 · 更新于 2026-09-24 · 约 25 分钟 · 12404 字 阅读 →
论文解读

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling

语音生成 | 7.5/10

 · 更新于 2026-09-24 · 约 18 分钟 · 8743 字 阅读 →
论文解读

X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning

语音克隆 | 8.0/10

 · 更新于 2026-09-24 · 约 17 分钟 · 8272 字 阅读 →
论文解读

JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions

音频质量评估 | 8.5/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6589 字 阅读 →
论文解读

Ecologically-Constrained Task Arithmetic for Multi-Taxa Bioacoustic Classifiers Without Shared Data

生物声学 | 8.0/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6468 字 阅读 →
论文解读

Learning Generalizable Action Representations via Pre-training AEMG

生物声学 | 7.5/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5844 字 阅读 →
论文解读

AC-Foley: Reference-Audio-Guided Video-to-Audio Synthesis with Acoustic Transfer

音频生成 | 8.0/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4618 字 阅读 →
论文解读

Alethia: A Foundational Encoder for Voice Deepfakes

音频深度伪造检测 | 8.0/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4189 字 阅读 →
论文解读

AVERE: Improving Audiovisual Emotion Reasoning with Preference Optimization

语音情感识别 | 8.0/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4850 字 阅读 →
论文解读

FlexiVoice: Enabling Flexible Style Control in Zero-Shot TTS with Natural Language Instructions

语音合成 | 8.0/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5075 字 阅读 →
论文解读

Learning multimodal dictionary decompositions with group-sparse autoencoders

跨模态检索 | 7.5/10

 · 更新于 2026-09-24 · 约 8 分钟 · 3971 字 阅读 →