To the line
Capabilities on the main line · ONE MULTIMODAL MODEL

Multimodal understanding

In one viewBy 2024, a single model could work jointly with text, images and audio instead of handing them between separate systems.

AI began perceiving a situation as a whole: seeing a screen, hearing speech and responding almost in real time. Robust understanding of long video and causal relationships remains ahead.

StatusPASSED
TypeCapabilities on the main line
Marker2024
Events in dossier8
Development chronology

Researched

2021-01

CLIP

About this eventA shared space for text and images — the foundation all image generation grew on.

CLIP learned to match images with internet text descriptions, placing two modalities in a shared representation space. It enabled visual categories to be specified with language instead of task-specific labels and became an important foundation for later generative and multimodal systems.

Source: OpenAI
2022-09

Whisper

About this eventSpeech recognition reaches the level where audio becomes an ordinary input.

Source: OpenAI
2023-04

Segment Anything

About this eventOne model learned to segment arbitrary objects from a point, box or prompt without training for each class.

Source: Meta AI

In progress

сейчас

Understanding long video

About this eventAn hour of footage fits, but following the plot and causes inside it does not work yet.

Distant horizons

впереди

Touch and smell

About this eventModalities with neither cheap sensors nor large datasets so far.

Sources and research

Primary material behind this dossier: papers, lab publications and official reports.

Directions
Capabilities on the main line