cd /news/artificial-intelligence/a-robot-duck-got-a-15mb-ai-brain · home › topics › artificial-intelligence › article
[ARTICLE · art-144585] src=stork.ai ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

A Robot Duck Got a 15MB AI Brain

Better Stack demonstrated a simulated Open Duck Mini robot duck controlled by Needle 3, a 15MB, 4-layer tool-calling model from Cactus Compute that maps natural-language requests to five actions: walk, turn, emote, dance, and shake head. Cactus Compute's Needle 3 family ranges from 8MB to 29MB and is built on a 121-million-parameter, 20-layer "intelligence ladder" architecture whose prefix slices of 2 to 20 layers each run as standalone models, with over half the parameters in an n-gram lookup table keeping practical compute near a 50-million-parameter model. The duck's manners were trained on roughly 1,700 generated examples plus 49 handwritten held-back prompts, and Cactus reports a 4-layer, 29-million-parameter version can outperform DeepSeek V4 Flash on specific tasks after fine-tuning.

by read5 min views1 publishedOct 3, 2026
A Robot Duck Got a 15MB AI Brain
Image: Stork (auto-discovered)

The duck obeys—but only if you say please #

A simulated robot duck walks, dances, or reacts only when a 15MB AI brain, Needle 3, maps a natural-language request to one of five available tools. This demonstration, from Better Stack, showcases a custom-tuned version of Cactus Compute’s compact foundation model operating a virtual Open Duck Mini, a physics-simulated model.

Needle 3 is not a chatbot; it does not converse or control the duck’s walking mechanics. Instead, a separate learned controller handles the duck’s movement, while Needle chooses the action (e.g., walk, turn, emote, dance, shake head). The model’s specific task is to dispatch these tools, even responding to nuances like requiring “please” for commands.

This focused design allows the model to be incredibly small and efficient. The 4-layer version used in the demo is only 15 megabytes, making it suitable for local execution on edge devices. This experiment highlights practical edge AI, where bounded tool dispatch can run offline, offering immediate, device-side responsiveness without cloud connectivity.

Cactus Compute’s Needle 3, part of their family of tool-calling models, ranges from 8 MB to 29 MB. Its "intelligence ladder" architecture allows developers to select a slice of the 121 million parameter model (with compute equivalent to 50M parameters) to fit diverse hardware, from phones to microcontrollers. This enables local AI capabilities in devices like the Pebble Index 01 Smart Ring.

One model, sliced to fit the device #

Needle 3 introduces an “intelligence ladder,” a 20-layer model designed for flexible deployment. Developers can take any prefix slice, from 2 to 20 layers, and each functions as a standalone model. This allows a single model download to serve diverse hardware, from tiny microcontrollers to more powerful smartphones.

Full Needle 3 features 121 million parameters, a significant jump from Needle 2’s 45 million. However, over half of these parameters reside in an n-gram lookup table. This design choice minimizes arithmetic operations, keeping the practical compute profile closer to a 50-million-parameter model, according to Cactus Compute.

This architecture creates a direct trade-off: larger slices offer enhanced capabilities, while smaller ones suit devices with tight hardware constraints. Needle is explicitly engineered as a tool-calling dispatcher, not for open-ended conversation. It excels at mapping natural language commands to a predefined set of actions, extracting parameters into structured JSON, or generating text embeddings for local search.

Its targeted function allows for remarkable efficiency. For instance, Cactus reports that a 4-layer version, with only 29 million parameters, can outperform DeepSeek V4 Flash on specific tasks after fine-tuning. This specialization enables Needle 3 to run on resource-constrained devices like the Pebble Index 01 Smart Ring, executing voice commands offline.

Teaching manners with 1,700 examples #

Teaching the robot duck manners required a clear set of rules for its five available tools: walk, turn, emote, dance, and shake head. The duck was trained to:

  • Obey requests containing "please."
  • Refuse bare commands by shaking its head.
  • Express anger at insults or threats, even if "please" was used.
  • React to dramatic events without needing "please."
  • Make no tool call for tasks it couldn't perform, like setting a timer.

To instill these behaviors, the creator generated approximately 1,700 training examples. Each example paired a natural language prompt with the desired tool response, ensuring the model learned underlying concepts rather than memorizing exact phrases. For instance, "pretend to be scared, please" mapped to a scared emote, while "I need you to remind me to buy milk" resulted in no call, as that function doesn't exist.

A separate set of 49 handwritten prompts was held back to test the model's generalization, ensuring it could apply its learned "manners" to novel situations. For more on the underlying technology, see Needle 3 - Foundation Model for Tiny Devices - Cactus Compute.

The fine-tuning process occurred on an RTX 5090 workstation. Training the full 20-layer Needle 3 model took about 12.5 minutes, while a smaller 4-layer version (approximately 15 megabytes) completed in about 5 minutes. A key caveat of this local fine-tuning was the loss of the model's confidence score, a feature typically available with Cactus Compute's hosted platform.

Enjoying this? Get one like it in your inbox each morning.

one email a day · unsubscribe in two clicks · no third-party tracking

The tiny model wins—by a narrower margin #

Held-out results clearly showed the benefit of fine-tuning. The baseline, untuned Needle 3 scored a meager 13 out of 49 on the test prompts. The fine-tuned, full 20-layer model achieved 41 out of 49, a significant improvement.

The roughly 15MB four-layer version scored 30 out of 49, a respectable performance for its size but still considerably lower than the full model. This smaller model sometimes misread an unrelated prompt, giving an odd reaction, or failed to identify the duck’s anger at insults. The performance depends heavily on the specific task and the quality of training examples.

These failures temper the story: while impressive, the small model’s abilities are not universal. Its success with the duck’s manners highlights its specialized design as a tool-calling dispatcher, not a general reasoning engine.

The real takeaway for edge AI is the potential for specialized, locally run dispatchers. They can handle specific, constrained tasks efficiently on small devices. However, benchmark claims and impressive demos should not be confused with broad reasoning ability or guaranteed product performance.

Frequently Asked Questions #

What is Needle 3?

Needle 3 is a compact model from Cactus Compute designed to dispatch structured tool calls on resource-constrained devices, rather than act as a general chatbot.

How can Needle 3 fit into a 15MB model?

Its 20-layer model can be sliced into smaller standalone versions. The duck demo used a fine-tuned four-layer version of about 15MB.

What did fine-tuning teach the robot duck?

It taught the duck to respond to polite requests, reject bare commands, react to rudeness or dramatic situations, and ignore requests beyond its abilities.

How well did the fine-tuned duck model perform?

On 49 handwritten test prompts, the four-layer model scored 30 correct, versus 13 for the untuned model and 41 for the fine-tuned 20-layer version.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @better stack 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/a-robot-duck-got-a-1…] indexed:0 read:5min 2026-10-03 · —