论文解读

Continuation Method for Feedback Delay Network Modal Decomposition

空间音频 | 6.5/10

 · 更新于 2026-09-11 · 约 8 分钟 · 3971 字 阅读 →
论文解读

Continuous-Token Diffusion for Speaker-Referenced TTS in Multimodal LLMs

语音合成 | 8.0/10

 · 更新于 2026-09-11 · 约 12 分钟 · 5822 字 阅读 →
论文解读

Contrastive Timbre Representations for Musical Instrument And Synthesizer Retrieval

音频检索 | 7.5/10

 · 更新于 2026-09-11 · 约 8 分钟 · 3791 字 阅读 →
论文解读

Controllable Embedding Transformation for Mood-Guided Music Retrieval

音乐检索 | 7.5/10

 · 更新于 2026-09-11 · 约 9 分钟 · 4452 字 阅读 →
论文解读

Cooperative Multi-Agent Reinforcement Learning for Adaptive Aggregation in Semi-Supervised Federated Learning with non-IID Data

联邦学习 | 7.0/10

 · 更新于 2026-09-11 · 约 9 分钟 · 4379 字 阅读 →
论文解读

CosyAccent: Duration-Controllable Accent Normalization using Source-Synthesis Training Data

语音转换 | 7.8/10

 · 更新于 2026-09-11 · 约 11 分钟 · 5037 字 阅读 →
论文解读

Coupling Acoustic Geometry and Visual Semantics for Robust Depth Estimation

空间音频 | 7.5/10

 · 更新于 2026-09-11 · 约 20 分钟 · 9555 字 阅读 →
论文解读

CoVA: Text-Guided Composed Video Retrieval for Audio-Visual Content

跨模态检索 | 6.5/10

 · 更新于 2026-09-11 · 约 9 分钟 · 4170 字 阅读 →
论文解读

Cross-Architecture Knowledge Distillation of WavLM for Lightweight Speaker Verification

说话人验证 | 8.0/10

 · 更新于 2026-09-11 · 约 13 分钟 · 6123 字 阅读 →
论文解读

Cross-Cultural Bias in Mel-Scale Representations: Evidence and Alternatives from Speech and Music

语音识别 | 7.0/10

 · 更新于 2026-09-11 · 约 9 分钟 · 4247 字 阅读 →
论文解读

Cross-Domain Contrastive Learning with Dynamic Threshold Calibration for Source Speaker Tracing

说话人验证 | 8.0/10

 · 更新于 2026-09-11 · 约 8 分钟 · 3746 字 阅读 →
论文解读

Cross-Lingual Alzheimer’s Disease Detection with Multimodal LLMs via Speech Cue-Augmented Prompting and Instruction Tuning

语音生物标志物 | 6.5/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4634 字 阅读 →
论文解读

Cross-Lingual F5-TTS: Towards Language-Agnostic Voice Cloning and Speech Synthesis

语音克隆 | 7.5/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4837 字 阅读 →
论文解读

Cross-Lingual Interleaving for Speech Language Models

语音大模型 | 7.5/10

 · 更新于 2026-09-11 · 约 11 分钟 · 5122 字 阅读 →
论文解读

Cross-Linguistic Rhythmic and Spectral Feature-Based Analysis of Nyishi and Adi: Two Under-Resourced Languages of Arunachal Pradesh

Cross-Linguistic Rhythmic and Spectral Feature-Based Analysis of Nyishi and Adi: Two Under-Resourced Languages of Arunachal Pradesh

 · 更新于 2026-09-11 · 约 1 分钟 · 34 字 阅读 →
论文解读

Cross-Modal Bottleneck Fusion for Noise Robust Audio-Visual Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-11 · 约 9 分钟 · 4237 字 阅读 →
论文解读

Cross-Modal Knowledge Distillation for Speech Large Language Models

语音大模型 | 7.0/10

 · 更新于 2026-09-11 · 约 8 分钟 · 3931 字 阅读 →
论文解读

CTC-DID: CTC-Based Arabic Dialect Identification for Streaming Applications

语音识别 | 6.5/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4764 字 阅读 →
论文解读

Curriculum Learning with Contrastive Loss for Lightweight Speaker Verification

说话人验证 | 6.5/10

 · 更新于 2026-09-11 · 约 10 分钟 · 4588 字 阅读 →
论文解读

Cutscene Agent: An LLM Agent Framework for Automated 3D Cutscene Generation

生成模型 | 8.5/10

 · 更新于 2026-09-11 · 约 12 分钟 · 5619 字 阅读 →