{"slug": "nobodywho-vs-cactus-on-device-inference-engine-comparison", "title": "NobodyWho vs Cactus: On-Device Inference Engine Comparison", "summary": "NobodyWho Edge runs GGUF models through llama.cpp with no conversion step, while Cactus uses its own Cactus Quants (CQ) rotation-and-codebook quantization format from 4-bit down to 1-bit, according to a technical comparison of the two on-device inference engines. Cactus hand-writes its CPU kernels in ARM NEON SIMD with Metal on Apple GPUs and the Apple Neural Engine for some vision models, while Qualcomm, MediaTek, and Exynos NPU support remains on its roadmap; NobodyWho uses Vulkan and Metal for GPU execution rather than targeting the NPU. NobodyWho installs via pip, flutter pub, npm, Maven Central, Swift Package Manager, and Godot's asset library, whereas Cactus requires cloning its repository and running a setup script, though it also publishes per-platform packages on pip, Maven, pub, and npm.", "body_md": "[← Back to blog](https://www.nobodywho.ai/posts/)\n\n# NobodyWho vs Cactus: On-Device Inference Engine Comparison\n\nChoosing an on-device inference engine sets where the model runs and under what terms. NobodyWho and Cactus both run models on the user's device, with no API key, no per-request cost, and no data leaving the hardware once the model is downloaded. Underneath, they are different engines. NobodyWho Edge is built on [llama.cpp](https://github.com/ggerganov/llama.cpp) and runs GGUF models. Cactus is a from-scratch engine with its own quantization format, and positions itself as a llama.cpp alternative. This is a technical comparison of NobodyWho vs Cactus across engine and model format, hardware, installation, platform coverage, cloud behaviour, and licensing. The better fit depends on the target device and the commercial model of the product.\n\n## **Engine and model format**\n\n[NobodyWho](https://github.com/nobodywho-ooo) runs GGUF models through llama.cpp. It loads any GGUF file from Hugging Face or a URL directly, with no conversion step, so any model already published in GGUF, at any of the quantization levels the ecosystem provides, runs as-is.\n\nCactus runs its own format. Cactus Quants (CQ) is a rotation-and-codebook quantization applied to every weight tensor, from 4-bit down to 1-bit, and the engine runs from CQ bundles. You get a bundle by downloading one Cactus has pre-built for its own model catalog (cactus download) or by converting a source model yourself (cactus convert), which [Cactus](https://github.com/cactus-compute/cactus) documents as experimental for models it has not pre-built.\n\nNobodyWho runs the existing GGUF catalog with no conversion step. Cactus's CQ is tuned in-house for on-device size and quality, with published accuracy tables per bit-width, at the cost of depending on its own bundles or an experimental conversion for anything outside its catalog.\n\nBoth also cover the parts of an app beyond chat: embeddings, speech-to-text, text-to-speech, and tool calling. NobodyWho generates the tool-calling grammar from your function signatures and constrains generation to it, so you pass plain functions and the output conforms to the expected structure without writing schemas by hand. Cactus takes OpenAI and MCP-style tool definitions, ships a vector index for retrieval, and provides Needle, a 26M-parameter model dedicated to tool calling.\n\n## **Hardware acceleration**\n\nCactus is built around the mobile processor. Its kernels are hand-written in ARM NEON SIMD for the CPU, with Metal on Apple GPUs and the Apple Neural Engine for some vision models. Qualcomm, MediaTek, and Exynos NPU support is on its roadmap, not shipped. Cactus publishes per-device benchmarks in its README, listing tokens per second and peak RAM for named iPhone and Mac hardware.\n\nNobodyWho takes a hardware-aware approach to acceleration, using Vulkan and Metal for GPU execution rather than targeting the NPU. This is the sharpest hardware split between them. On older or low-end phones that lean on the CPU, Cactus's hand-written kernels are the advantage. On devices with a capable GPU, NobodyWho runs on the accelerator most platforms already expose.\n\n## **Installation and first call**\n\nNobodyWho installs through each ecosystem's own package manager: *pip install nobodywho* for Python, *flutter pub add nobodywho* for Flutter, *npm install react-native-nobodywho* for React Native, *ai.nobodywho:nobodywho* from Maven Central for Kotlin, Swift Package Manager for Swift, and the in-editor asset library for Godot. Each binding is a thin wrapper over one shared Rust core.\n\nCactus installs its engine by cloning the repository and running a setup script, then building the binding and downloading a model bundle:\n\n```\ngit clone https://github.com/cactus-compute/cactus && cd cactus && source ./setup \ncactus build --python \ncactus download LiquidAI/LFM2-VL-450M\n```\n\nCactus also publishes per-platform packages on pip, Maven, pub, and npm.\n\nThe difference carries into the first call. NobodyWho's Python API handles model loading and lifecycle for you, and returns the reply as a string:\n\n``` python\nfrom nobodywho import Chat\n\nchat = Chat(\"hf:NobodyWho/Qwen_Qwen3-0.6B-GGUF/Qwen_Qwen3-0.6B-Q4_K_M.gguf\") \nresponse = chat.ask(\"What is the capital of Denmark?\").completed() \nprint(response)\n```\n\nCactus's Python binding is a ctypes FFI over its C engine, with explicit model lifecycle and JSON message payloads:\n\n``` python\nfrom cactus import ensure_model, cactus_init, cactus_complete, cactus_destroy\nimport json\n \nbundle = ensure_model(\"LiquidAI/LFM2-VL-450M\")\nmodel = cactus_init(str(bundle), None, False)\nmessages = json.dumps([{\"role\": \"user\", \"content\": \"What is the capital of Denmark?\"}])\nresult = cactus_complete(model, messages, None, None, None)\nprint(result[\"response\"])\ncactus_destroy(model)\n```\n\nBoth snippets are each project's own documented quick-start.\n\n## **Platform support**\n\nCactus targets mobile and embedded. Its C and C++ core runs on wearables, smart-home devices, robots, and Raspberry Pi, with bindings for Flutter, React Native, Kotlin Multiplatform, Swift, Python, and Rust.\n\nNobodyWho targets mobile, desktop (Linux, macOS, Windows), Python, and the JVM, and ships a binding for the [Godot game engine](https://docs.nobodywho.ooo/godot/).\n\nBoth engines cover phones and wearables, and Cactus reaches further into embedded hardware like robots and Raspberry Pi. Cactus also supports desktop deployment on macOS and ARM Linux. NobodyWho ships a Godot binding and targets desktop app runtimes across Linux, macOS, and Windows through its JVM and Python bindings. Cactus has no game-engine binding. Neither ships a browser or WebAssembly target. NobodyWho has an open GitHub issue tracking WASM export, and Cactus does not target the browser.\n\n## **Cloud behaviour**\n\nBoth run inference locally by default. Cactus adds an optional cloud handoff that routes low-confidence queries to a hosted model. NobodyWho has no cloud router in the engine. In deployments that prohibit outbound network calls, NobodyWho has nothing to disable, and Cactus's handoff must be turned off and verified.\n\n## **Licensing**\n\nNobodyWho uses [EUPL-1.2](https://github.com/nobodywho-ooo/nobodywho), an OSI-approved open-source licence. It permits proprietary and commercial use with no revenue limit. The only obligation is that redistributing a modified version of the engine requires publishing those engine changes.\n\nCactus is source-available under its own licence. That [licence](https://github.com/cactus-compute/cactus/blob/main/LICENSE) grants free use to individuals, students, non-profits, and organizations with \"Less than $2,000,000 USD in total funding\" and \"Less than $2,000,000 USD in gross annual revenue.\" An organization that does not meet those criteria \"must obtain a separate commercial license,\" and a qualifying organization that later crosses either threshold has \"thirty (30) days\" to obtain one. Both licences were verified on 9 September 2026. Cactus has revised its terms before, so check the current file before shipping a commercial product.\n\nBoth engines are free under those thresholds. Once a company crosses either one, Cactus requires a paid commercial licence and NobodyWho stays free.\n\n## **Choosing between NobodyWho and Cactus**\n\nThe decision follows target hardware and commercial model.\n\n| Requirement | Engine | \n|---|---|\n| Run any existing GGUF model with no conversion step | NobodyWho | \n| In-house quantization tuned for on-device size and quality | Cactus | \n| Hand-optimized CPU performance on ARM without a capable GPU | Cactus | \n| Robotics, smart-home, or Raspberry Pi | Cactus | \n| Desktop applications alongside mobile | NobodyWho | \n| A model running locally inside a Godot game | NobodyWho | \n| OSI open-source licence with no revenue ceiling | NobodyWho | \n\nThe two are different engines with different model formats, so the choice is set by where the model has to run and how the licence treats the product, not by a shared core. The [NobodyWho source and per-binding docs](https://github.com/nobodywho-ooo/nobodywho) list current platform and format support.", "url": "https://wpnews.pro/news/nobodywho-vs-cactus-on-device-inference-engine-comparison", "canonical_source": "https://www.nobodywho.ai/posts/nobodywho-vs-cactus/", "published_at": "2026-09-16 00:00:00+00:00", "updated_at": "2026-09-30 09:16:31.353494+00:00", "lang": "en", "topics": ["ai-infrastructure", "ai-tools", "large-language-models", "developer-tools"], "entities": ["NobodyWho", "NobodyWho Edge", "Cactus", "llama.cpp", "Cactus Quants", "GGUF", "Hugging Face", "Needle"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/nobodywho-vs-cactus-on-device-inference-engine-comparison", "markdown": "https://wpnews.pro/news/nobodywho-vs-cactus-on-device-inference-engine-comparison.md", "text": "https://wpnews.pro/news/nobodywho-vs-cactus-on-device-inference-engine-comparison.txt", "jsonld": "https://wpnews.pro/news/nobodywho-vs-cactus-on-device-inference-engine-comparison.jsonld"}}