Standard Intelligence bets raw screen video can train general computer agents better than language-based tool workflows
EDITOR BRIEF
Standard Intelligence is pursuing a contrarian approach to AI agents: training models directly on raw screen recordings of computer use rather than text, screenshots, and tool calls. Its system predicts mouse movements, clicks, and keystrokes from pixels, backed by an 11-million-hour computer action dataset and a highly token-efficient video encoder.
INSIGHTS
The strategy reflects a broader push toward scalable, data-driven agent training inspired by the “bitter lesson”: less hand-engineering, more raw compute and data. If successful, video pre-training could become a major alternative to today’s language-model-centric agent architectures for automating knowledge work.
COMMENTS
Discussion
> geekhaus:~$ next read?

