GEEK HAUS
Back to feed

Standard Intelligence bets raw screen video can train general computer agents better than language-based tool workflows

·Sequoia Capital
read original

EDITOR BRIEF

Standard Intelligence is pursuing a contrarian approach to AI agents: training models directly on raw screen recordings of computer use rather than text, screenshots, and tool calls. Its system predicts mouse movements, clicks, and keystrokes from pixels, backed by an 11-million-hour computer action dataset and a highly token-efficient video encoder.

INSIGHTS

The strategy reflects a broader push toward scalable, data-driven agent training inspired by the “bitter lesson”: less hand-engineering, more raw compute and data. If successful, video pre-training could become a major alternative to today’s language-model-centric agent architectures for automating knowledge work.

COMMENTS

Discussion

> geekhaus:~$ next read?

Next read recommendations