Qwen 3.8-Max and Claude Opus 5 show why raw benchmark scores don't predict the bill

EDITOR BRIEF
Alibaba promoted Qwen 3.8-Max as near the top of coding-agent benchmarks, but an independent VulcanBench run found it mid-pack at best and last by default. The article argues both views can be valid because Alibaba allowed far longer run times, showing that token and time budgets heavily shape model performance and cost.
INSIGHTS
For reasoning models, price per token is becoming a weak proxy for real-world expense because failed attempts and hidden thinking tokens can dominate the bill. Buyers should evaluate cost per successful task and make time or token limits explicit before choosing models for production workflows.
COMMENTS
Discussion
> geekhaus:~$ next read?
Next read recommendations

The Verge
Voters mostly don’t like AI and data centers, but neither party seems to have an edge
TechCrunch
The AI data center boom is colliding with cities scarred by big industry
rheinmetall.github.io