Stop adding more GPUs: Weka's new storage platform reduces load by caching 100% of an AI model's pre-calculated tokens

EDITOR BRIEF
Weka unveiled NeuralMesh 6 and its self-designed Wekapod 3 hardware, using aggregated NAND flash to extend GPU memory for production AI workloads. The platform targets long-context inference by caching pre-calculated tokens, reducing repeated computation and freeing GPU capacity for more users or responses.
INSIGHTS
The launch reflects a broader shift in AI infrastructure from simply adding GPUs to improving GPU utilization with storage-layer optimization. If enterprises can reliably offload memory pressure to cheaper flash, inference costs may fall and AI deployments could scale faster despite constrained GPU supply.
COMMENTS
Discussion
> geekhaus:~$ next read?
Next read recommendations

VentureBeat
Google’s Gemini 3.8 Flash is built for agents, while its Cyber twin hunts vulnerabilities

VentureBeat
Meta prices Muse Voice Transcribe at $0.18 an hour, with real-time diarization for 20+ speakers: a steal for enterprises?

VentureBeat