CLIP learned to match images with internet text descriptions, placing two modalities in a shared representation space. It enabled visual categories to be specified with language instead of task-specific labels and became an important foundation for later generative and multimodal systems.
Source: OpenAICLIP
About this eventA shared space for text and images — the foundation all image generation grew on.