
Ke Li 🍁
279 posts

Ke Li 🍁
@KL_Div
Assistant Professor and Canada CIFAR AI Chair @SFU @AmiiThinks. Ph.D. from @Berkeley_EECS and HBSc from @UofTCompSci. Formerly @GoogleAI and Member of @the_IAS.



There have been some follow-up questions on the relationship between XMs and IMLE, so I’m going to provide some clarifications (some of this was in E.3 of the paper). IMLE is a great paper we respect and cite in XM, and I highly recommend people check it out. In fact, I’d encourage people to read the two papers side by side. I think this would help clarify the different motivations and let the merits of each speak for themselves. That being said, XMs are a generalization of IMLE and of best-of-K methods more broadly, and the recent claim that “XM is just a special case of IMLE” is inaccurate. XMs are about one thing, which is increasing generative expressivity by factoring the training loop, and IMLE is one specific instance of that (end-to-end Forward XMs). Because of this, >90% of the results in the XM paper go against the IMLE theory, and have not been studied by IMLE (including the entire 3rd pretraining axis portion of the paper, which was done by combining XMs with existing scalable reconstructive generative models, which IMLE has never focused on). Additionally, the XM paper's main contribution is empirical insights on generative expressivity, not best-of-K. In the paper, we explicitly state, “We do not claim to invent best-of-K”, and we also cite IMLE within the first two paragraphs of the approach section and in other locations. 🧵Thread:

We discovered a third pretraining axis beyond parameters and data: exploration. Scaling exploration monotonically improves existing models across images/video/language, and unlocks end-to-end generation. In the simplest case, it's just a for loop. Introducing Explorative Modeling. TLDR: - Gains from exploration grow with scale: 7%→36% as data scales, 13%→23% as parameters scale, and gains double at 3× the compute - Adding exploration to ~SOTA baselines improves data efficiency by 6.2×, FLOP efficiency by 4.1×, parameter efficiency by 47%, and hits a near-SOTA 1.43 unguided FID on ImageNet - Exploration lets you trade training compute for generalization, and scales how end-to-end your generative model is - End-to-end Explorative Models (XMs) match diffusion performance on control tasks with up to 256× less inference compute 🧵Thread:




