GEEK HAUS
Back to feed
2026/07/27/humanlayer-tests-claude-opus-5-on-slopcodebench

HumanLayer tests Claude Opus 5 on SlopCodeBench, finding gains on long-horizon coding but still low pass rates

·github.com
read original

EDITOR BRIEF

HumanLayer benchmarked Claude Opus 5, Opus 4.8, and Sonnet 5 on a subset of SlopCodeBench, a long-horizon coding benchmark that reveals requirements gradually across checkpoints. Opus 5 led the group with a 24% strict pass rate on the tested subset, but the author says overall performance remains weak.

INSIGHTS

SlopCodeBench highlights a gap between coding agents that can solve isolated tasks and those that can maintain code quality as requirements evolve. The low scores suggest long-horizon software engineering remains a major frontier for AI coding tools, even as frontier models improve.

COMMENTS

Discussion

> geekhaus:~$ next read?

Next read recommendations