Publications

For the latest citation information, please visit my Google Scholar profile.

H3-World: Turning Language Understanding into World Control

Technical Report

Danze Chen, Zeqing Wang, Ziyue Lin, Xingyi Yang, Yeying Jin

Turns the language understanding already learned by a large video generator into precise, temporally grounded character and camera control with lightweight adaptation.

Supervise What Survives: Geometry-Guided VLA Adaptation from Synthetic Robot Videos

CoRL 2026

Danze Chen, Yanzhe Chen, Qiming Huang, Zhijun Cao, Chen Gao, Mike Zheng Shou

Introduces geometry-guided representation alignment for adapting vision-language-action models from synthetic robot videos while keeping low-level control grounded in real demonstrations.

WorldMind: Decoupled Game World Model for State-Aware NPC Behavior

CCF-A Conference Under Review

Zhiyang Deng, Boran Zhang, Danze Chen, Yeying Jin

Decouples state understanding, NPC decision-making, action control, and video generation, then reconnects them in a closed interaction loop for state-aware NPC behavior.

ReactiveGWM: Steering NPC in Reactive Game World Models

CCF-A Conference Under Review

Zeqing Wang, Danze Chen, Zhaohu Xing, Zizhao Tong, Yinhan Zhang, Xingyi Yang, Yeying Jin

Separates low-level player control from high-level NPC strategy, enabling controllable interactions and zero-shot strategy transfer across game world models.

LayerTracer text-to-SVG generation and layer-wise vectorization examples

LayerTracer: Cognitive-Aligned Layered SVG Synthesis via Diffusion Transformer

ICCV 2025 Oral

Yiren Song, Danze Chen, Mike Zheng Shou

Learns cognitively aligned layer-by-layer SVG construction with a diffusion transformer, producing editable vector graphics with meaningful semantic structure.

Notation: My name is shown in bold.