Milestones on the line · SENSES

Multimodality

You can now show the same model a picture, play it a recording or hand it a video — and it understands them together rather than separately.

A single model takes text, image, audio and video and answers in real time. The computer got senses instead of a text box.

StatusPASSED
TypeMilestones on the line
Marker2024
What exists today6

Researched

2021

CLIP

A shared space for text and images — the foundation all image generation grew on.

verified
2022

Whisper

Speech recognition reaches the level where audio becomes an ordinary input.

verified
2023

GPT-4V

Vision as a standard part of a frontier model.

verified
2024

Realtime voice

Conversation latency drops to human level.

verified

In progress

сейчас

Understanding long video

An hour of footage fits, but following the plot and causes inside it does not work yet.

NEW

Distant horizons

впереди

Touch and smell

Modalities with neither cheap sensors nor large datasets so far.

NEW
Branches
Milestones on the line
To the line

Track the line as it moves

Once a week: which branches advanced, what unlocked, and what turned out to be overstated.