Cactus unveils Needle 2, a 14MB agentic LLM designed to run on phones, wearables, smart homes, and robots
EDITOR BRIEF
Cactus released Needle 2, a 14MB agentic LLM for tool calling, device control, and structured extraction on low-power devices. The model runs a full session in 28MB of RAM and claims high decode speeds on Raspberry Pi 5, VR headsets, and budget Android phones while competing with much larger small models on mobile-use benchmarks.
INSIGHTS
Needle 2 reflects a shift in edge AI from desktop-class hardware toward ultra-small models that can run continuously on cheap, battery-limited devices. If its benchmark claims translate to real-world reliability, it could make local assistants more practical for IoT, wearables, smart homes, and consumer robots without relying on cloud inference.
COMMENTS
Discussion
> geekhaus:~$ next read?
Next read recommendations
cactuscompute.com
Whistle debuts as a 16.9MB on-device speech recognition model for CPUs, supporting seven languages and fast tool-call workflows
nishtahir.com
Guide shows how to build a fast LLM decision model by constraining outputs to fixed answer choices

The Verge