Predicting Video in Representation Space

A working plan: encode each video frame with a frozen image transformer, compress it into a handful of continuous visual tokens, and train a causal transformer to predict the next frame’s tokens — learning how scenes evolve without ever translating them into text or pixels.

September 28, 2026 · 5 min