Predicting Video in Representation Space
A working plan: encode each video frame with a frozen image transformer, compress it into a handful of continuous visual tokens, and train a causal transformer to predict the next frame’s tokens — learning how scenes evolve without ever translating them into text or pixels.