DuplexGen: Decoupling Content, Timing, and Acoustics for Synthetic Dialogue Speech
📄 DuplexGen: Decoupling Content, Timing, and Acoustics for Synthetic Dialogue Speech 标签:#语音合成 #语音大模型 #语音识别 #说话人日志 #数据集 6.1/10 | 创新 1.3/2 | 严谨 1/1.5 | 实验 1/1.5 | 清晰 0.7/1 | 影响 0.9/1.5 | 开源 0/1.5 | 复现 0.1/0.5 | 工程 1.1/1.5 ✅ 6.1/10 | 前50% | 文档类型:方法研究 | 评分置信度:中 | #语音合成 | #语音大模型 | #语音识别 #说话人日志 | arxiv 👥 作者与机构 第一作者:Pengcheng Wang(Department of Information and Communications Engineering, Institute of Science Tokyo, Yokohama, Japan) 通讯作者:未说明 作者列表:Pengcheng Wang(Department of Information and Communications Engineering, Institute of Science Tokyo, Yokohama, Japan)、Sheng Li(Department of Information and Communications Engineering, Institute of Science Tokyo, Yokohama, Japan)、Jiyi Li(Hokkaido University, Hokkaido, Japan)、Takahiro Shinozaki(Department of Information and Communications Engineering, Institute of Science Tokyo, Yokohama, Japan) 💡 毒舌点评 亮点是把对话内容、交互时序与声学渲染拆成三个责任边界明确的阶段,并用 script-constrained duplex decoding 把“说什么”锁死、“何时说”留给两个互听模型,确实在 FTO 分布上显著优于两个不会产生 overlap 的拼接基线。但缺点同样明显:所谓 emergent timing 仍被 handoff window、max_padding、bc_refractory 三个旋钮和“仅限话轮转换区”的重叠约束限制,宏观动态不是完全涌现;论文又没有公开任何代码、权重或数据集访问方式,声明的 released benchmark 无法被审稿人核查。综合看,是一篇中规中矩、证据覆盖面不错但离完整资源贡献仍有距离的工作。 ...