NobodyWho vs RunAnywhere: On-Device Inference Engine Comparison A side-by-side test of on-device LLM inference engines NobodyWho and RunAnywhere found RunAnywhere slows with each conversational turn, reaching up to 33× slower after 20 turns on Android, while NobodyWho starts answering sooner on a single prompt. The comparison, run on an iPhone Air and a Samsung S25 using the React Native SDKs, also found RunAnywhere accepts only one image per request with no audio and forgets previous tool results, whereas NobodyWho mixes several images and audio files in one prompt. NobodyWho is licensed under EUPL-1.2 and stays free regardless of company funding or revenue, while RunAnywhere's Apache 2.0-based license requires a paid commercial license once an organization exceeds $1,000,000 USD in total funding or $1,000,000 USD in gross annual revenue. ← Back to blog https://www.nobodywho.ai/posts/ NobodyWho vs RunAnywhere: On-Device Inference Engine Comparison On paper, NobodyWho and RunAnywhere look almost identical. They run LLMs locally on consumer devices such as laptops and phones, are built on llama.cpp, and both list the same features across Kotlin, Swift, Python, Flutter and React Native. The differences only show up once you put them in the same app and start a real conversation. That's exactly what I did, side by side on iPhone and Android. Here is a brief summary of the technical findings: - Speed: NobodyWho starts answering sooner on a single prompt, and RunAnywhere gets slower with every turn of a conversation, up to 33× slower after 20 turns on Android. - Multimodal: RunAnywhere accepts one image per request and no audio. NobodyWho lets you mix several images and audio files in one prompt. - Tool calling: RunAnywhere forgets previous tool results, so it can't answer follow-up questions about them. Every result can be reproduced with my test app on GitHub https://github.com/pielouNW/runanywhere-react-native-starter-app . Before getting to the benchmarks, let's start with an overview of both libraries: engine, model format and features, platform support and licensing. Engine, model format and features NobodyWho https://github.com/nobodywho-ooo/nobodywho runs any GGUF model through llama.cpp, loaded straight from Hugging Face, a URL or a local path, with no conversion step. RunAnywhere https://github.com/RunanywhereAI/runanywhere-sdks also uses llama.cpp, and can additionally register MLX and QHexRT Qualcomm Hexagon NPU backends. Registering them did not change any of my results below. Both libraries offer hardware acceleration and a similar feature set: text generation, multimodal input, embeddings, RAG, speech-to-text, text-to-speech, structured output, voice activity detection and tool calling. Platform support NobodyWho and RunAnywhere both support Kotlin, Swift, Python, Flutter, React Native and Expo. RunAnywhere also supports Electron and WebAssembly. NobodyWho has started work https://github.com/nobodywho-ooo/nobodywho/pull/755 on WebAssembly, which will be available soon. NobodyWho also ships for Godot, and runs on Apple Vision Pro https://apps.apple.com/us/app/nobodywho-eyes/id6771770762 and Apple Watch https://apps.apple.com/us/app/nobodywho-wrist/id6762020355?platform=watch . Licensing NobodyWho uses EUPL-1.2 https://github.com/nobodywho-ooo/nobodywho/blob/main/LICENSE , an OSI-approved open-source licence, allowing proprietary and commercial projects to use the engine free of charge. However, if you distribute a modified version of NobodyWho itself, those engine changes must be open sourced. RunAnywhere describes its licence as "RunAnywhere License Apache 2.0 based, with additional commercial-use terms ." https://github.com/RunanywhereAI/runanywhere-sdks/blob/main/LICENSE For companies, the free grant only applies to organizations with both "Less than $1,000,000 USD in total funding" and "Less than $1,000,000 USD in gross annual revenue." Organizations outside the listed criteria "must obtain a separate commercial license." If either threshold is exceeded later on, "a commercial license must be obtained within thirty 30 days." In practice, NobodyWho stays free no matter how much funding your company raises or how much revenue it earns. RunAnywhere is free for a company only while it stays under $1M in total funding and $1M in annual revenue. Once it passes either figure, a paid commercial licence is required. Technical comparison To compare both engines under the same conditions, I added NobodyWho to the RunAnywhere React Native Starter App https://github.com/RunanywhereAI/react-native-starter-app and ran both side by side on an iPhone Air and a Samsung S25. The comparison covers speed , multimodal input and tool calling . I used the React Native SDKs, but none of these issues are specific to React Native. They come from RunAnywhere's inference engine and APIs, so they affect every platform and language it supports. You can check out the test app https://github.com/pielouNW/runanywhere-react-native-starter-app and run it on your own device to reproduce every result. Note that the RunAnywhere starter app needed a few fixes before it would build with Xcode 27, which you can find listed in the appendix at the end of this article. Speed The speed tests were done on both phones with the Qwen3 0.6B model and the same configuration. On a single prompt, NobodyWho starts answering sooner on both phones, but RunAnywhere generates faster on Android: - iPhone Air: RunAnywhere generates 70.5 tokens per second tok/s , with a time to first token TTFT of 201 ms. NobodyWho generates 70.8 tok/s, with a TTFT of 46 ms. - Samsung S25: RunAnywhere generates 52.4 tok/s, with a TTFT of 1268 ms fluctuate between 300 and 1300 ms in my tests . NobodyWho generates 32.2 tok/s, with a TTFT of 230 ms. The real problem appears in a conversation, where RunAnywhere's TTFT grows with every turn . Over 20 turns, it climbs from 147 ms to 471 ms on the iPhone Air 3× slower , and from 333 ms to over 11 seconds on the S25 33× slower . NobodyWho stays between 33 ms and 60 ms on the iPhone Air, and between 218 ms and 896 ms on the S25. It means that after a few prompts, NobodyWho is answering faster , even on Android. This happens because on every turn, RunAnywhere starts from an empty cache and re-processes the system prompt, the entire chat history and the new message. NobodyWho keeps the conversation in its KV cache and only processes the new message. With RunAnywhere, the longer the conversation, the slower the response . NobodyWho currently shows the same growing TTFT issue with hybrid models like Qwen3.5, since the cache can't yet be reused between turns in the same way. A fix https://github.com/nobodywho-ooo/nobodywho/pull/803 is in progress. Multimodal input Multimodal LLMs like Gemma 4 can take images and/or audio files along with a prompt. Let's see how each library handles them. With RunAnywhere, you cannot send an audio file to the model, and you can only send one image at a time : js // node modules/@runanywhere/core/src/Public/Api/Vlm.ts export const vlm = { async generate image: ImageInput, prompt: string, options?: LlmOptions : Promise