Struggling Toward Generative Model Evaluation
A blog post on why generative model evaluation is hard, how representation-based metrics became the default language, and why benchmarks are useful but never final.
Research notes, project writeups, and observations on generative modeling.
A blog post on why generative model evaluation is hard, how representation-based metrics became the default language, and why benchmarks are useful but never final.
We benchmark EMA decay across pixel-, latent-, and representation-space diffusion models and show that the popular 0.9999 default silently trades recall for precision.
Dual-Stage Registers address outlier tokens in both the encoder and diffusion transformer, improving ImageNet-256 FID from 5.89 to 4.58 at 80 epochs.
A unified f-divergence framework connecting more than 10 one-step diffusion distillation methods, with a one-step FID of 1.02 on ImageNet 64 × 64. NeurIPS 2025.