Web AI runtimes: five ways to run a language model in the browser Nearform built a demo web application that runs five in-browser AI runtimes side by side through a uniform chat, info and configuration interface, with source published at github.com/nearform/web-ai-demo. The company reported that browser inference remains constrained: context windows in the tested setups range from about 1K to 9K tokens, with Chrome's Prompt API reporting a contextWindow of 9,216 tokens on its test machine, and model downloads range from roughly 180 MB to 19 GB. On an iPhone 15 Pro, Nearform also compared how the five runtimes handled tool calling. Your new AI platform is the browser At Nearform, we build all manner of AI applications and solutions for our clients. While most folks think of AI models backing applications like these as solely the purview of the backend, like an API server routing to a cloud-based frontier model, there’s a whole emerging world of potential on the frontend, in the browser. We’ve already written about https://nearform.com/digital-community/browser-based-vector-search-fast-private-and-no-backend-required/ how things like semantic vector search can be performant, private, and versatile in the browser, and this same innovation wave applies to AI models as well. Web browsers now have direct access to the GPU processing and memory needed for AI inference tasks thanks to WebGPU https://developer.mozilla.org/en-US/docs/Web/API/WebGPU API , an API standard gaining increasing support https://caniuse.com/webgpu in modern browsers. AI models have also made enormous gains in reducing model size while improving inference quality. This makes for an exciting time: we can run real AI inference in the browser, with plenty of options. And we do mean a lot of options. This area is so new and hot that there are a growing number of in-browser AI frameworks to choose from. So, we built a demo web application to see which ones actually hold up, running five of the most promising options side by side in a uniform chat, info, and configuration interface. You can follow along, gory details and all, at https://nearform.github.io/web-ai-demo/ https://nearform.github.io/web-ai-demo/ with source available at https://github.com/nearform/web-ai-demo https://github.com/nearform/web-ai-demo . Most of what follows — the tips, the tricks, and the stumbling blocks along the way — came out of getting five very different runtimes to answer the same question through the same interface. After the runtime tours, we’ll also get into two comparisons we found especially interesting: how the runtimes handled tool calling, and what happened when we put all five on an iPhone 15 Pro. … but browser inference is pretty challenging OK, so it’s both exciting … and hard. In-browser inference models have real constraints worth reviewing at a high level, as these concerns will thread throughout our tour of the various runtimes. - Context is really small. Consumers of frontier and open-weight models are used to context in the 250K to 1M token range. Many of the model setups we’ll tour here run from about 1K up to 9K total context. That’s tens to hundreds of times smaller and often barely enough to fit a reasonable system prompt, let alone domain-specific context information. Chrome’s Prompt API, for example, reported a contextWindow of 9,216 tokens on our test machine. Others e.g., those supporting full Gemma 4 can be set to modern ranges of 128K–256K, though on our M5 Max, answers from the larger Gemma 4 models became garbled past about 32K. Browser models are mostly comparable to models from the 2022–2023 era of AI, which is to say behind today, yet a capable start. But this is changing quickly, with Gemma 4 signalling a new wave of models that are qualitatively much closer to modern frontier models. - The model downloads are large. To get a model on-device, you’re looking at a download of anywhere from roughly 180 MB up to a whopping 19 GB. That will noticeably impact the UI. Fortunately, all of our browser inference runtimes have a means of caching downloaded models. Additionally, there are open source cache libraries https://github.com/jasonmayes/web-ai-model-proxy-cache you can use now, as well as an early W3C WICG Cross-Origin Storage https://wicg.github.io/cross-origin-storage/ proposal that would let you download an AI model once in a browser and then make it available to all the different apps that might use it. Transformers.js https://github.com/huggingface/transformers.js/blob/836b9cb1b3542f217eab330f6fc62768bdf314f9/packages/transformers/src/utils/cache/CrossOriginStorageCache.js , wllama https://github.com/ngxson/wllama/issues/246 , and WebLLM https://github.com/mlc-ai/web-llm/pull/748 already support it. If you’re impatient, there are already multiple browser extensions https://github.com/web-ai-community/cross-origin-storage-extension available-on that let you use it now - Memory is limited, particularly on iOS. On a desktop browser, a runtime that streams the model in rather than loading it all at once can use close to the GPU’s full memory — our M5 Max ran every Gemma 4 model Google publishes for LiteRT-LM.js, up to the 19 GB 31B. Many recent Android phones have room for Gemma 4-class models too. iOS is the exception: it caps how much memory a browser tab can take https://bugs.webkit.org/show bug.cgi?id=268816 c10 , and since every iOS browser runs on WebKit, Chrome for iOS gets the same limit as Safari. Go over it and there’s no error to catch, as iOS just kills the tab. - The small models are, well, small. Some models fit within iOS’s limits, but they’re small: the ones that ran on our iPhone were all around 400 MB or smaller. At the current state of the art, it’s a real challenge to produce acceptable results for most real-world use cases. That means you’re going to get not just hallucinations, but often gibberish or nonsensical responses. You can limit the impact of this with better base context, but that pushes against the other problem of extremely limited context. The clearest example we saw is one model at two quantisations. Qwen3.5-0.8B at Q4 K M 533 MB answered coherently; the same model at Q2 K XL 418 MB , right at the edge of what our iPhone could load, often produced gibberish responses and missed tool calls in our testing. - Everything is changing all the time. The AI technology ecosystem moves really fast. The web AI ecosystem, we’d argue, moves even faster . By the time this article is published and you read it, the capabilities, APIs, and limitations of much of what we’ll talk about will have changed — and potentially dramatically. For anything you end up trying out and you should try this stuff out , make sure to head over to the project docs first and check for updates on the latest advice on getting started. More constraints will turn up as you go. We’ll cover a few more in this article and the demo, but just be aware that it’s a wild web AI world out there. The five, and how we compared them The five we built into the demo are the ones we found most compelling, and we’ll walk you through the most basic starting point for each, along with the comparison notes from running them side by side. Our examples all circle around fairly basic questions that a model should be able to answer out of the box, with no additional context, plus a simple follow-up — consider this the “hello world” of web AI. The online demo defaults to a one-shot “In one sentence, what is bikeshedding?”, which we’ll use as a running check on output quality at every stop. The code examples here take an easier, multi-turn “What does a web browser do?” question, for a slightly different flavour if you’re running all five yourself. You can try the snippets on your own, or follow along in the demo app https://nearform.github.io/web-ai-demo/?runtime=chrome-prompt-api , which adds diagnostics, control levers, and supporting information about each runtime in various panes and drill-downs. Demo app showing wllama runtime inference with Gemma 4 model. And to lead in with the takeaways for all the runtimes, we’ll kick things off with a brief comparison table: | Runtime | What it is | Reach for it when | |---|---|---| | Chrome Prompt API https://developer.chrome.com/docs/ai/prompt-api | Gemini Nano https://developer.chrome.com/docs/ai/built-in in Chrome or Phi-4-mini in Edge, shipped inside the browser | Your users are on Chrome or Edge https://learn.microsoft.com/en-us/microsoft-edge/web-platform/prompt-api preview desktop, and you want good answers easily | | WebLLM https://github.com/mlc-ai/web-llm | MLC https://llm.mlc.ai/ -compiled weights on WebGPU | You want a predictable, pre-compiled catalogue and an OpenAI-compatible API | | wllama https://github.com/ngxson/wllama | llama.cpp https://github.com/ggml-org/llama.cpp compiled to WASM; runs GGUFs from Hugging Face | You want to keep up with Hugging Face as new GGUFs land | | Transformers.js https://huggingface.co/docs/transformers.js | Hugging Face’s transformers API in JavaScript, running ONNX weights on ONNX Runtime Web https://onnxruntime.ai/docs/tutorials/web/ | You already live in the Transformers ecosystem and want its ONNX catalogue | | LiteRT-LM.js https://developers.google.com/edge/litert-lm/js | The web build of Google’s LiteRT-LM runtime the same engine as on Android and iOS , taking over from MediaPipe’s LLM Inference API; loads -web.litertlm / -gpu.litertlm bundles | You want Gemma 4, from E2B up to 31B, tuned for GPU speed | That table says what each runtime is. The difference you’ll feel first is who owns the conversation. Two of the five — the Chrome Prompt API and LiteRT-LM.js — keep it themselves, so each turn is just a prompt, and both also take the system prompt at construction rather than per turn so editing it mid-conversation does nothing . The other three — WebLLM, wllama, and Transformers.js — leave it to you: you manage the message array and resend it every turn. You’ll see that split in the differing ask signatures in every snippet that follows. Another significant difference is tool calling, which we cover after the individual runtime tours. Chrome Prompt API — promising, if you’re on Chrome Our first runtime is already available in modern Chrome browsers https://caniuse.com/wf-languagemodel , is part of an evolving W3C standard https://webmachinelearning.github.io/prompt-api/ , and is heading towards broader browser and model support. The Prompt API https://developer.chrome.com/docs/ai/prompt-api provides a conversational AI model using Gemini Nano that accepts multimodal inputs https://developer.chrome.com/docs/ai/prompt-api add expected input and output and generates text outputs. The size of the downloaded Gemini Nano model may vary by browser — you can find more info at the special chrome://on-device-internals URL in Chrome. It is also one of a family of task-specific AI APIs https://developer.chrome.com/docs/ai/built-in in Chrome, alongside the Translator https://developer.chrome.com/docs/ai/translator-api , Summarizer https://developer.chrome.com/docs/ai/summarizer-api , and Proofreader https://developer.chrome.com/docs/ai/proofreader-api APIs. Let’s see the minimum code you need to get up and running with the Prompt API: if await LanguageModel.availability === "unavailable" { throw new Error "No built-in model on this device" ; } const session = await LanguageModel.create { initialPrompts: { role: "system", content: "Answer in one sentence." } , } ; // The session keeps the history, so a turn is just a prompt. const ask = async prompt = { let output = ""; for await const chunk of session.promptStreaming prompt { output += chunk; } return output; }; console.log await ask "What does a web browser do?" ; console.log await ask "Now make it shorter." ; session.destroy ; Pretty straightforward We wrap the model call in ask to standardise our example code across the five runtimes. Head over to the live demo https://nearform.github.io/web-ai-demo/?runtime=chrome-prompt-api in a modern Chrome browser and try out the API and the various settings for yourself. On the bikeshedding question, we’ve seen typically good responses like: Bikeshedding is debating trivial details that don't significantly impact the overall outcome or decision. As we progress through our other runtimes, we’ll see that answering this correctly with essentially no context is a non-trivial bar. The current Gemini Nano model produces very good outputs. Although context is technically “dynamic”, we usually see 9,200+ tokens of available context, which is pretty spacious for the web AI environment. Being limited to Chrome and Edge, in preview at present limits where the runtime runs, but it has a caching advantage at least until Cross-Origin Storage lands : once you enable and download an initial model, it’s available for all apps using the Prompt API. There’s a lot more on the roadmap for the Prompt and related APIs, and you can get a sneak peek by joining the Built-in AI: Early Preview Program https://developer.chrome.com/docs/ai/join-epp . Our takeaway for the Prompt API is that if you’re in a Chrome environment where it’s supported, and you’re OK with a proprietary Gemini model, then this is a solid API to build around. The future looks a bit more open, with the open-licensed Gemma 4 model landing behind a flag chrome://flags/ gemma4-for-built-in-ai , see source https://chromium-review.googlesource.com/c/chromium/src/+/7808999 in Chrome Canary v153+ . WebLLM — a curated catalogue on WebGPU WebLLM https://webllm.mlc.ai/ is a long-standing well, for this arena project that runs MLC https://mlc.ai/ -compiled models with WebGPU support. It has a curated catalogue of known working models, which is good for predictability, though new models can take a while to arrive, and there’s a backlog of open model requests https://github.com/mlc-ai/web-llm/issues?q=is%3Aissue%20is%3Aopen%20%22model%20request%22 . Its newest released families are Gemma 3 and Qwen3.5, but some requests are moving: Gemma 4 support text and audio has landed on the main branch https://github.com/mlc-ai/web-llm/pull/793 , with prebuilt catalogue entries due in the next release. You aren’t limited to the catalogue, though: you can compile your own models for architectures MLC already supports now including Gemma 4 E2B https://github.com/mlc-ai/mlc-llm/pull/3559 with the MLC LLM toolchain https://llm.mlc.ai/docs/deploy/webllm.html . It also supports very small models that fit within iOS’s tight memory limits — one of only two runtimes we got an answer out of on our test iPhone without crashing, which we’ll come back to later. And while the models have reasonable inference speed, the catalogue’s default context is small — 1K, 2K, or 4K, chosen to save memory. You can raise it per model if the device has memory to spare once we did, Qwen2.5-0.5B found a code we buried 16K tokens deep . WebLLM provides a basic OpenAI-compatible API. Let’s see what hooking it up looks like for a multi-turn conversation. js import { CreateMLCEngine } from "https://cdn.jsdelivr.net/npm/@mlc-ai/web-llm@0.2.85/+esm"; if await navigator.gpu?.requestAdapter { throw new Error "web-llm needs WebGPU" ; } // Downloaded on first use, then cached in the browser. const engine = await CreateMLCEngine "SmolLM2-360M-Instruct-q4f16 1-MLC", { initProgressCallback: { text } = console.log text , } ; // Stateless, OpenAI-style: resend the whole conversation each turn. // The engine reuses its KV cache, so only the new turn is prefilled. const ask = async messages = { let output = ""; for await const chunk of await engine.chat.completions.create { messages, stream: true, } { output += chunk.choices 0 ?.delta?.content ?? ""; } return output; }; const messages = { role: "system", content: "Answer in one sentence." }, { role: "user", content: "What does a web browser do?" }, ; const first = await ask messages ; console.log first ; // Turn two: keep our own history, then resend all of it. messages.push { role: "assistant", content: first } ; messages.push { role: "user", content: "Now make it shorter." } ; console.log await ask messages ; await engine.unload ; As with any runtime, output quality depends mostly on the model, and the small end of the catalogue struggles. You can see for yourself in our WebLLM demo https://nearform.github.io/web-ai-demo/?runtime=web-llm . The smallest option, SmolLM2-135M-Instruct-q0f16-MLC , makes a good stress test: we put the bikeshedding question to it four separate times. It never got it right, and it was never wrong in the same way twice: That's all, Millie. Cloth-style bikeshedding is a style of pulling a bike onto a bike wheel, tagging the wheels with stains or fans to colour them, and rearranging them into a frame for a new bike. Think about it: jumping forward from present things and caring about the future. Is this enough information to answer the user's question? The second response is arguably the only one that’s not pure nonsense or a non-answer. Oh, and we hope Millie’s doing OK. Moving up the catalogue helps. The far larger Hermes-3-Llama-3.1-8B-q4f16 1-MLC https://nearform.github.io/web-ai-demo/?runtime=web-llm&model=Hermes-3-Llama-3.1-8B-q4f16 1-MLC gets it right, though it ignored “in one sentence” and went on for a full paragraph, plus a second