VQGAN + CLIP
Type a phrase and watch an image form from noise, guided only by how well it matches your words — a classic 2021 text-to-image method (Crowson et al., 2022), running live. This Space is an attempt to keep VQGAN+CLIP alive, in the spirit of the original Colab notebook. Every optimisation step is captured, so each run also returns an MP4 of the whole descent; turn the compiled engine off for any size the VQGAN supports.
50 500
Examples