DeepSWE blows up the AI coding leaderboard, crowns GPT-5.5, and finds Claude Opus exploiting a benchmark loophole

EDITOR BRIEF
Datacurve released DeepSWE, a 113-task AI coding benchmark across 91 open-source repositories and five languages, showing a much wider performance gap among frontier models than SWE-Bench Pro. OpenAI’s GPT-5.5 led with a 70% score, 16 points ahead of the nearest rival, while Datacurve says SWE-Bench Pro graders produced incorrect verdicts in about one-third of reviewed trials.
INSIGHTS
The results suggest enterprise buyers may need more realistic and auditable coding evaluations before committing to AI developer tools. If the claimed benchmark error rate is validated, it could weaken confidence in leaderboard-driven marketing and push the industry toward tougher, more transparent evaluations.
COMMENTS
Discussion
> geekhaus:~$ next read?
Next read recommendations

VentureBeat
Google’s Gemini 3.8 Flash is built for agents, while its Cyber twin hunts vulnerabilities

VentureBeat
Meta prices Muse Voice Transcribe at $0.18 an hour, with real-time diarization for 20+ speakers: a steal for enterprises?

VentureBeat