Nearby in the stack

Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models · arXivDesk