2026/08/07/deepseek-v4-flash-0731-posts-strong-arc-agi
DeepSeek V4 Flash 0731 posts strong ARC-AGI benchmark scores with low per-task costs across three reasoning levels
EDITOR BRIEF
DeepSeek says its V4 Flash 0731 model reaches 89.0% on ARC-AGI-1 Semi-Private and 61.4% on ARC-AGI-2 Semi-Private at maximum reasoning effort. The company also reports high and low reasoning variants, showing a performance-cost tradeoff across the ARC-AGI evaluations.
INSIGHTS
The results suggest frontier-style reasoning performance is becoming cheaper to run, which could pressure rivals to publish both accuracy and cost metrics. The three-tier setup reflects a broader trend toward configurable inference, letting users choose between speed, price, and reasoning depth.
COMMENTS
Discussion
> geekhaus:~$ next read?

