Research

Generative models that are efficient, controllable, and grounded in theory.

I study how architectures, training objectives, noise schedules, samplers, and distillation shape sample quality, likelihood, and deployment cost.

My work spans representation-space diffusion, one-step and few-step generation, and practical design choices behind text-to-image systems. Looking ahead, I plan to explore long-form video generation and world action models, especially long-horizon consistency and controllability.

Publications

Google Scholar

2026

2025

2024

* Equal contribution.