GEEK HAUS
Back to feed

DeepSWE blows up the AI coding leaderboard, crowns GPT-5.5, and finds Claude Opus exploiting a benchmark loophole

·VentureBeat
read original
DeepSWE blows up the AI coding leaderboard, crowns GPT-5.5, and finds Claude Opus exploiting a benchmark loophole

EDITOR BRIEF

Datacurve released DeepSWE, a 113-task AI coding benchmark across 91 open-source repositories and five languages, showing a much wider performance gap among frontier models than SWE-Bench Pro. OpenAI’s GPT-5.5 led with a 70% score, 16 points ahead of the nearest rival, while Datacurve says SWE-Bench Pro graders produced incorrect verdicts in about one-third of reviewed trials.

INSIGHTS

The results suggest enterprise buyers may need more realistic and auditable coding evaluations before committing to AI developer tools. If the claimed benchmark error rate is validated, it could weaken confidence in leaderboard-driven marketing and push the industry toward tougher, more transparent evaluations.

COMMENTS

Discussion

> geekhaus:~$ next read?

Next read recommendations