At Waymo, an AI project isn't ready until its evals are — not when the model performs well

EDITOR BRIEF
Waymo engineering leader Manasi Joshi said the company uses eval-centric development to decide when autonomous driving AI is ready, making testing a core engineering discipline rather than a final prelaunch step. The company applies continuous evaluations during training, after training, and in simulations to manage high-stakes real-world risk.
INSIGHTS
Waymo’s approach highlights a broader enterprise lesson: AI systems need measurable reliability before deployment, especially as companies adopt agents for customer service, coding, finance, and operations. The trend points toward evaluation infrastructure becoming as important as model performance in enterprise AI governance.
COMMENTS
Discussion
> geekhaus:~$ next read?
Next read recommendations
virtualizationhowto.com
Broadcom removes public VMware VDDK downloads, potentially complicating migrations from vSphere to rival virtualization platforms

The Verge
Audi’s new A2 E-tron is its most affordable and efficient EV yet
TechCrunch