论文解读

FUSEMOS: Perceptual Evaluation of Text-to-Music Generation with Dual-Encoder Fusion and Ranking-Aware Composite Loss

音乐生成 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4701 字 阅读 →
论文解读

Game-Time: Evaluating Temporal Dynamics in Spoken Language Models

语音对话系统 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4572 字 阅读 →
论文解读

Hashing-Baseline: Rethinking Hashing in the Age of Pretrained Models

音频检索 音频分类 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4362 字 阅读 →
论文解读

HD-PPT: Hierarchical Decoding of Content- and Prompt-Preference Tokens for Instruction-Based TTS

语音合成 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4633 字 阅读 →
论文解读

How to Label Resynthesized Audio: The Dual Role of Neural Audio Codecs in Audio Deepfake Detection

音频深度伪造检测 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4091 字 阅读 →
会议任务专题

ICASSP 2026 - 模型评估

共 16 篇 ICASSP 2026 模型评估 方向论文

 · 更新于 2026-09-25 · 约 51 分钟 · 25201 字 阅读 →
论文解读

Identifying the Minimal and Maximal Phonetic Subspace of Speech Representations

语音识别 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4910 字 阅读 →
论文解读

Identity Leakage Through Accent Cues in Voice Anonymisation

语音匿名化 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4208 字 阅读 →
论文解读

Improving the Speaker Anonymization Evaluation’s Robustness to Target Speakers with Adversarial Learning

语音匿名化 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5692 字 阅读 →
论文解读

Integrating Speaker Embeddings and LLM-Derived Semantic Representations for Streaming Speaker Diarization

说话人分离 | 6.5/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8207 字 阅读 →
论文解读

Interpretable Music Harmonic Analysis Through Multilinear Mixture of Experts

音乐理解 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4568 字 阅读 →
论文解读

Investigating Modality Contribution in Audio LLMs for Music

模型评估 | 6.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3688 字 阅读 →
论文解读

Investigating The Effect Of Sentence-Level Syntactic Structure On Information Loss In The Human Auditory System

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3548 字 阅读 →
论文解读

Learning Linearity in Audio Consistency Autoencoders via Implicit Regularization

音频生成 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4229 字 阅读 →
论文解读

Leveraging Large Speech Language Models as Evaluators for Expressive Speech

语音情感识别 | 6.5/10

 · 更新于 2026-09-25 · 约 7 分钟 · 3233 字 阅读 →
论文解读

Leveraging Multiple Speech Enhancers for Non-Intrusive Intelligibility Prediction for Hearing-Impaired Listeners

模型评估 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5540 字 阅读 →
论文解读

Leveraging prediction entropy for Automatic prompt weighting in Zero-Shot Audio-Language Classification

音频分类 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4393 字 阅读 →
论文解读

Lingometer: On-Device Personal Speech Word Counting System

语音活动检测 | 8.0/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5825 字 阅读 →
论文解读

LLAC: Learned Lossless Audio Codec

音频无损编码 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3743 字 阅读 →
论文解读

Measuring Prosody Diversity in Zero-Shot TTS: A New Metric, Benchmark, and Exploration

语音合成 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4416 字 阅读 →