Turn Text into Actions with Needle, 14Mb
Cactus Compute released Needle 2, a 14MB function-calling LLM that runs entirely on a Raspberry Pi 5's CPU and selected the correct tool for the prompt "Turn the LED on" in 78 milliseconds. The model,…
Cactus Compute released Needle 2, a 14MB function-calling LLM that runs entirely on a Raspberry Pi 5's CPU and selected the correct tool for the prompt "Turn the LED on" in 78 milliseconds. The model,…
Cactus Compute released Needle 3, an open-weight foundation model that ships as a single 8-29MB file for on-device tool calling, structured extraction, and embeddings, with a runtime engine under 1MB …
Cactus Compute released Needle 3, a single-file on-device foundation model that ships between 8 and 29 MB depending on layer depth and handles tool calling, structured extraction and text embeddings w…
Cactus Compute's Needle 3, a foundation model shipping as a single 8-29 MB file, beats models ten times its size on mobile tool-calling accuracy and matches models two to three times larger on structu…
Cactus Compute released Needle 3 on September 17, an Apache 2.0-licensed automation foundation model that ships as an 8 to 29MB binary and runs tool calling, structured extraction, and text embeddings…
Cactus Compute released Needle 2, a 14 MB agentic model designed for tool calling, device control, and structured data extraction on edge devices. A developer's hands-on testing revealed several bugs,…
Cactus Compute released Needle 2, an open 45M-parameter tool-calling model that ships as a 14MB binary and runs a full session in about 28MB of RAM, targeting devices with no GPU or NPU. The model ach…
Cactus Compute released Needle2, a 45-million-parameter language model compressed to a 14MB binary that runs AI agent tool-calling at 500 tokens per second on a Raspberry Pi 5 and 6,000 tokens per sec…
Cactus Compute's Needle 2 model fits a functional LLM into just 14MB, using Simple Attention Networks to reduce per-token computation to 70 MFLOPs, compared to 164 MFLOPs for a standard transformer of…
Cactus Compute, a Y Combinator Summer 2025 startup, released Needle 2, a 14MB language model for tool calling on low-memory devices, capable of running on phones, wearables, home devices, robots, and …