·arcprize.org
DeepSeek V4 Flash 0731 posts strong ARC-AGI benchmark scores with low per-task costs across three reasoning levels
DeepSeek says its V4 Flash 0731 model reaches 89.0% on ARC-AGI-1 Semi-Private and 61.4% on ARC-AGI-2 Semi-Private at maximum reasoning effort. The company also reports high and low...
read →