{"slug": "nobodywho-vs-runanywhere-on-device-inference-engine-comparison", "title": "NobodyWho vs RunAnywhere: On-Device Inference Engine Comparison", "summary": "A side-by-side test of on-device LLM inference engines NobodyWho and RunAnywhere found RunAnywhere slows with each conversational turn, reaching up to 33× slower after 20 turns on Android, while NobodyWho starts answering sooner on a single prompt. The comparison, run on an iPhone Air and a Samsung S25 using the React Native SDKs, also found RunAnywhere accepts only one image per request with no audio and forgets previous tool results, whereas NobodyWho mixes several images and audio files in one prompt. NobodyWho is licensed under EUPL-1.2 and stays free regardless of company funding or revenue, while RunAnywhere's Apache 2.0-based license requires a paid commercial license once an organization exceeds $1,000,000 USD in total funding or $1,000,000 USD in gross annual revenue.", "body_md": "[← Back to blog](https://www.nobodywho.ai/posts/)\n\n# NobodyWho vs RunAnywhere: On-Device Inference Engine Comparison\n\nOn paper, NobodyWho and RunAnywhere look almost identical. They run LLMs locally on consumer devices such as laptops and phones, are built on llama.cpp, and both list the same features across Kotlin, Swift, Python, Flutter and React Native. The differences only show up once you put them in the same app and start a real conversation. That's exactly what I did, side by side on iPhone and Android.\n\nHere is a brief summary of the technical findings:\n\n- **Speed:** NobodyWho starts answering sooner on a single prompt, and RunAnywhere gets slower with every turn of a conversation, up to 33× slower after 20 turns on Android.\n- **Multimodal:** RunAnywhere accepts one image per request and no audio. NobodyWho lets you mix several images and audio files in one prompt.\n- **Tool calling:** RunAnywhere forgets previous tool results, so it can't answer follow-up questions about them.\n\nEvery result can be reproduced with my [test app on GitHub](https://github.com/pielouNW/runanywhere-react-native-starter-app).\n\nBefore getting to the benchmarks, let's start with an overview of both libraries: engine, model format and features, platform support and licensing.\n\n## Engine, model format and features\n\n[NobodyWho](https://github.com/nobodywho-ooo/nobodywho) runs any GGUF model through llama.cpp, loaded straight from Hugging Face, a URL or a local path, with no conversion step.\n\n[RunAnywhere](https://github.com/RunanywhereAI/runanywhere-sdks) also uses llama.cpp, and can additionally register MLX and QHexRT (Qualcomm Hexagon NPU) backends. Registering them did not change any of my results below.\n\nBoth libraries offer hardware acceleration and a similar feature set: text generation, multimodal input, embeddings, RAG, speech-to-text, text-to-speech, structured output, voice activity detection and tool calling.\n\n## Platform support\n\nNobodyWho and RunAnywhere both support Kotlin, Swift, Python, Flutter, React Native and Expo.\n\nRunAnywhere also supports Electron and WebAssembly. NobodyWho has [started work](https://github.com/nobodywho-ooo/nobodywho/pull/755) on WebAssembly, which will be available soon. NobodyWho also ships for Godot, and runs on [Apple Vision Pro](https://apps.apple.com/us/app/nobodywho-eyes/id6771770762) and [Apple Watch](https://apps.apple.com/us/app/nobodywho-wrist/id6762020355?platform=watch).\n\n## Licensing\n\nNobodyWho uses [EUPL-1.2](https://github.com/nobodywho-ooo/nobodywho/blob/main/LICENSE), an OSI-approved open-source licence, allowing proprietary and commercial projects to use the engine free of charge. However, if you distribute a modified version of NobodyWho itself, those engine changes must be open sourced.\n\nRunAnywhere describes its licence as [\"RunAnywhere License (Apache 2.0 based, with additional commercial-use terms).\"](https://github.com/RunanywhereAI/runanywhere-sdks/blob/main/LICENSE) For companies, the free grant only applies to organizations with both \"Less than $1,000,000 USD in total funding\" and \"Less than $1,000,000 USD in gross annual revenue.\" Organizations outside the listed criteria \"must obtain a separate commercial license.\" If either threshold is exceeded later on, \"a commercial license must be obtained within thirty (30) days.\"\n\nIn practice, NobodyWho stays free no matter how much funding your company raises or how much revenue it earns. RunAnywhere is free for a company only while it stays under $1M in total funding and $1M in annual revenue. Once it passes either figure, a paid commercial licence is required.\n\n## Technical comparison\n\nTo compare both engines under the same conditions, I added NobodyWho to the [RunAnywhere React Native Starter App](https://github.com/RunanywhereAI/react-native-starter-app) and ran both side by side on an iPhone Air and a Samsung S25. The comparison covers **speed**, **multimodal input** and **tool calling**.\n\nI used the React Native SDKs, but none of these issues are specific to React Native. They come from RunAnywhere's inference engine and APIs, so they affect every platform and language it supports. You can [check out the test app](https://github.com/pielouNW/runanywhere-react-native-starter-app) and run it on your own device to reproduce every result. Note that the RunAnywhere starter app needed a few fixes before it would build with Xcode 27, which you can find listed in the appendix at the end of this article.\n\n### Speed\n\nThe speed tests were done on both phones with the Qwen3 0.6B model and the same configuration.\n\nOn a single prompt, NobodyWho starts answering sooner on both phones, but RunAnywhere generates faster on Android:\n\n- **iPhone Air:** RunAnywhere generates 70.5 tokens per second (tok/s), with a time to first token (TTFT) of 201 ms. NobodyWho generates 70.8 tok/s, with a TTFT of 46 ms.\n- **Samsung S25:** RunAnywhere generates 52.4 tok/s, with a TTFT of 1268 ms (fluctuate between 300 and 1300 ms in my tests). NobodyWho generates 32.2 tok/s, with a TTFT of 230 ms.\n\nThe real problem appears in a conversation, where **RunAnywhere's TTFT grows with every turn**. Over 20 turns, it climbs from 147 ms to 471 ms on the iPhone Air (3× slower), and from 333 ms to over 11 seconds on the S25 (33× slower). NobodyWho stays between 33 ms and 60 ms on the iPhone Air, and between 218 ms and 896 ms on the S25. It means that after a few prompts, **NobodyWho is answering faster**, even on Android.\n\nThis happens because on every turn, RunAnywhere starts from an empty cache and re-processes the system prompt, the entire chat history and the new message. NobodyWho keeps the conversation in its KV cache and only processes the new message. With RunAnywhere, **the longer the conversation, the slower the response**.\n\nNobodyWho currently shows the same growing TTFT issue with hybrid models like Qwen3.5, since the cache can't yet be reused between turns in the same way. A [fix](https://github.com/nobodywho-ooo/nobodywho/pull/803) is in progress.\n\n### Multimodal input\n\nMultimodal LLMs like Gemma 4 can take images and/or audio files along with a prompt. Let's see how each library handles them.\n\nWith RunAnywhere, you cannot send an audio file to the model, and you can only send **one image at a time**:\n\n``` js\n// node_modules/@runanywhere/core/src/Public/Api/Vlm.ts\n\nexport const vlm = {\n  async generate(\n    image: ImageInput,\n    prompt: string,\n    options?: LlmOptions\n  ): Promise<GenerationResult> { ... },\n\n  generateStream(\n    image: ImageInput,\n    prompt: string,\n    options?: LlmOptions\n  ): AsyncIterable<GenerationEvent> { ... },\n};\n```\n\nIn both `generate` and `generateStream`, the image is a required argument. Every follow-up question about the same image means sending it again and having the model process it again, instead of continuing the conversation naturally.\n\nIt is also **not possible to interleave** several images and audio files in a single prompt, which NobodyWho allows:\n\n``` js\nconst response = await chat\n  .ask(\n    new Prompt([\n      Prompt.Text(\"Tell me what you see in the images and what you hear in the audio.\"),\n      Prompt.Image(\"/path/to/dog.png\"),\n      Prompt.Image(\"/path/to/cat.png\"),\n      Prompt.Audio(\"/path/to/sound.mp3\"),\n    ]),\n  )\n  .completed();\n```\n\n### Tool calling\n\nTool calling lets your LLM call predefined functions when needed. For example, if you give the LLM a `get_weather` function, it will call it whenever the user asks about the weather.\n\nBoth libraries handle tool calling properly, but RunAnywhere does not keep previous tool calls and their results in the conversation history, so follow-up questions get answered incorrectly.\n\nThe screenshot above compares the official NobodyWho and RunAnywhere apps, both running Qwen3 4B. After a `get_weather` call, NobodyWho answers \"What is the humidity?\" from the earlier result, while RunAnywhere has forgotten it and asks for the location again.\n\n## Choosing between NobodyWho and RunAnywhere\n\nThe two libraries share the same feature list, but the differences become obvious once you use them. With RunAnywhere, the chat slows down with every message, the assistant forgets what its tools just returned, and you can only send one image at a time.\n\n| Requirement | Engine | \n|---|---|\n| OSI open-source licence with no revenue ceiling | NobodyWho | \n| Wide platform and language support | Both | \n| Fast multi-turn conversations | NobodyWho | \n| Multimodal support | NobodyWho | \n| Tool calling | NobodyWho | \n\nRunAnywhere is a good fit if you need Electron or WebAssembly support right now. For everything else, **NobodyWho is the better choice for on-device AI**: answers stay fast as conversations grow, tool results carry over to follow-up questions, and the licence stays free however big your company gets.\n\n## Appendix: building the RunAnywhere starter app with Xcode 27\n\nThe RunAnywhere starter app doesn't build out of the box with Xcode 27 on macOS 27. Both the starter app and the latest RunAnywhere SDK use older versions of React Native ([0.83](https://github.com/RunanywhereAI/react-native-starter-app/blob/e1117fe0e506f1d5edbb148f0d179b75b7f6c7b7/package.json#L25) and [0.85](https://github.com/RunanywhereAI/runanywhere-sdks/blob/acc341c8eae9078a5ab99102bad0ca8bb0377fc7/bindings/react-native/package.json#L63)) instead of the current [0.87](https://reactnative.dev/versions), and they conflict with Xcode 27's stricter compiler. To get it running, I [fixed four issues in the Podfile](https://github.com/pielouNW/runanywhere-react-native-starter-app/blob/f70c06dc344792f9320ee42dad6974dda5c69126/ios/Podfile#L44-L90) (an outdated fmt C++ library, Swift 6 concurrency errors, an Objective-C class hidden from Swift and an undeclared static library), and a method called by the wrong name with a [Yarn patch](https://github.com/pielouNW/runanywhere-react-native-starter-app/blob/f70c06dc344792f9320ee42dad6974dda5c69126/.yarn/patches/@runanywhere-core-npm-0.20.19-f4d947d688.patch).\n\n*Info*\n\n- This article was written on October 9, 2026. Results may have changed since then, on both the NobodyWho and RunAnywhere sides.\n- The `@runanywhere` dependencies in`package.json` are set to v0.20.19.[Updating them](https://github.com/pielouNW/runanywhere-react-native-starter-app/tree/feat/ra-v0.20.27) to the latest published version, v0.20.27, did not change the results.", "url": "https://wpnews.pro/news/nobodywho-vs-runanywhere-on-device-inference-engine-comparison", "canonical_source": "https://www.nobodywho.ai/posts/nobodywho-vs-runanywhere/", "published_at": "2026-10-09 00:00:00+00:00", "updated_at": "2026-10-09 08:23:00.889990+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "ai-tools", "developer-tools"], "entities": ["NobodyWho", "RunAnywhere", "llama.cpp", "React Native", "iPhone Air", "Samsung S25", "EUPL-1.2", "Qualcomm Hexagon NPU"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/nobodywho-vs-runanywhere-on-device-inference-engine-comparison", "markdown": "https://wpnews.pro/news/nobodywho-vs-runanywhere-on-device-inference-engine-comparison.md", "text": "https://wpnews.pro/news/nobodywho-vs-runanywhere-on-device-inference-engine-comparison.txt", "jsonld": "https://wpnews.pro/news/nobodywho-vs-runanywhere-on-device-inference-engine-comparison.jsonld"}}