GEEK HAUS
Back to feed
2026/08/06/qwen-3-8-max-and-claude-opus-5-show-why-raw

Qwen 3.8-Max and Claude Opus 5 show why raw benchmark scores don't predict the bill

·VentureBeat
read original
Qwen 3.8-Max and Claude Opus 5 show why raw benchmark scores don't predict the bill

EDITOR BRIEF

Alibaba promoted Qwen 3.8-Max as near the top of coding-agent benchmarks, but an independent VulcanBench run found it mid-pack at best and last by default. The article argues both views can be valid because Alibaba allowed far longer run times, showing that token and time budgets heavily shape model performance and cost.

INSIGHTS

For reasoning models, price per token is becoming a weak proxy for real-world expense because failed attempts and hidden thinking tokens can dominate the bill. Buyers should evaluate cost per successful task and make time or token limits explicit before choosing models for production workflows.

COMMENTS

Discussion

> geekhaus:~$ next read?

Next read recommendations