architectureIntroduced by VQ-VAE · 2017
Discrete visual tokens (VQ-VAE)
Compress images into a vocabulary of discrete codes so Transformers can model them like text.
Drafted by AI · not yet reviewed
How this idea evolved
No earlier ideas recorded for this concept yet.