GENERATIVE MODELS / VISUAL INTELLIGENCE
Building generative models that understand structure and create with precision.
I am Haoyang Tong. I recently completed my bachelor's degree at the School of Artificial Intelligence, University of Chinese Academy of Sciences. In September 2026, I will begin the joint PhD program between ZGCA and CASIA, advised by Prof. Ran He in MAIS & NLPR, CASIA.
My research explores generative models, large language models, and world models, with a current focus on controllable generation across pixels and 3D space.
LATEST
News
Our paper Energy-Guided Flow Matching is now available on arXiv.
CoGrad3D was accepted by AAAI 2026.
BACKGROUND
Education
CASIA & ZGCA
Joint PhD Program · Advisor: Prof. Ran He
MAIS & NLPR, Institute of Automation, Chinese Academy of SciencesUniversity of Chinese Academy of Sciences
Bachelor's Degree · Artificial Intelligence
School of Artificial IntelligenceRESEARCH & INDUSTRY
Experience
JD.com
Research Intern
Pixel-level generative modeling · Platform Product & R&D CenterMAIS & NLPR, CASIA
Research Intern · Advisor: Prof. Ran He
3D generation with 2D diffusion priorsSELECTED WORK
Publications
TEXT-TO-3D GENERATION
CoGrad3D: Spatially-Coupled Timestep Optimization with Orthogonal Gradient Fusion for 3D Generation
Improving geometric consistency and texture fidelity through adaptive sampling and cross-view gradient fusion.
PIXEL-SPACE GENERATION
Energy-Guided Flow Matching
A coarse-to-fine flow-matching trajectory guided by image-specific spectral energy.
PIXEL-SPACE GENERATION
Pixel-Space Diffusion via Observation Operators
Time-dependent observation operators align diffusion supervision and decoder refinement with coarse-to-fine image recovery.
IMAGE SEGMENTATION
iFAN: Inference-Aware Learning for Plain Mask Transformers
Quality-aware query ranking and cross-layer self-distillation improve mask transformers while retaining efficient final-layer inference.
3D & 4D GENERATION
Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation
Multimodal language models detect and guide corrections of spatial and temporal inconsistencies in 3D and 4D generation.
LAYERED IMAGE GENERATION
LiWi: Layering in the Wild
High-fidelity natural image decomposition with agent-driven data synthesis, shadow-aware learning, and boundary correction.



