{"slug": "webml-community-runs-three-qwen3-5-models-locally-in-browsers-with-webgpu", "title": "Webml-community runs three Qwen3.5 models locally in browsers with WebGPU", "summary": "Webml-community released a WebGPU demo that runs Alibaba's Qwen3.5 Small models (0.8B, 2B, and 4B) locally in browsers, using Transformers.js for the JavaScript runtime and WebGPU for hardware acceleration. The demo, highlighted by Hugging Face engineer Joshua Lochner, allows users to chat with the models without a cloud API, keeping data on-device. It excludes the 9B variant, underscoring the practical limits of browser-based deployment.", "body_md": "Qwen3.5 Small shipped on March 2, 2026, according to [Qwen's Qwen3.5 release announcement](https://qwen.ai/blog?id=qwen3.5&ref=runtimewire). A WebGPU demo now shows how its 0.8B, 2B and 4B models can run locally in a browser. The demo is maintained by webml-community and was highlighted by [Joshua Lochner (@xenovacom)](https://x.com/xenovacom?ref=runtimewire), the Hugging Face engineer associated with Transformers.js. The browser demonstration is a separate deployment project from Alibaba's Qwen3.5 release announcement.\n\n[Joshua Lochner's Qwen3.5 browser demonstration on X](https://twitter.com/xenovacom/status/2089435071384076306?ref=runtimewire)\n\nThe [Qwen3.5 WebGPU demo](https://huggingface.co/spaces/webml-community/Qwen3.5-WebGPU?ref=runtimewire) gives Alibaba's five-and-a-half-month-old small-model family a straightforward deployment path: open a webpage, choose a model, attach an image if needed and start a conversation. The inference runs inside the browser without a cloud API, according to the demo, and chats remain on the user's device.\n\nThe demo turns [Qwen3.5 Small's downloadable weights](https://github.com/QwenLM/Qwen3.5?ref=runtimewire) into a browser-based test. [Transformers.js](https://github.com/huggingface/transformers.js?ref=runtimewire) provides the JavaScript runtime layer, while WebGPU puts compatible models on a user's own graphics hardware. That lowers the barrier between finding a model and trying it without setting up a server or paying for inference.\n\nRuntimeWire's August 17 Qwen3.8 report covered [newly released model weights and license terms](/article/alibaba-qwen3-8-open-weights-27b-local-ai). This demo concerns a separate deployment path for an older model family: webml-community packages three Qwen3.5 Small models for local browser inference through Transformers.js and WebGPU.\n\n### A browser is the deployment target\n\nThe demo offers Qwen3.5 models at 0.8B, 2B and 4B parameters. Users can submit text or images, and the models support reasoning-oriented tasks.\n\nThose details make the demo more consequential than a hosted chatbot skin. Local inference can keep private inputs off an external API, remove per-request charges and continue without a permanent connection after the required files have been downloaded. It also shifts the compute expense from a model provider to the user's hardware, which can appeal to application developers serving frequent, narrowly defined tasks.\n\nThe browser demo has a clear boundary. Alibaba released four Qwen3.5 Small models on March 2, 2026: 0.8B, 2B, 4B and 9B. The webml-community interface stops at 4B. It does not state a minimum memory configuration, provide generation speeds across different laptops or promise that every WebGPU-capable device will deliver the same experience.\n\nThe omission of the 9B version is instructive. Model weights can be downloadable while still being inconvenient for a browser tab. The local-model market is increasingly decided by the complete deployment path: quantization, runtime support, memory use, kernel performance and a usable interface. Parameter count alone tells developers little about whether a model will fit comfortably into an actual product.\n\nThe [Transformers.js repository](https://github.com/huggingface/transformers.js?ref=runtimewire) supports language, vision, audio and other model classes without requiring a server. For Qwen, that work turns downloadable weights into something a developer can test before opening a terminal.\n\n### The frontier claim needs a machine attached\n\nAlibaba's post repeats a frontier-performance claim without identifying the model, laptop hardware or benchmark behind it. The browser demo supports the narrower claim that the smaller models can run locally through WebGPU.\n\n[Qwen / Alibaba on X](https://x.com/Alibaba_Qwen/status/2089597078385446969?ref=runtimewire)\n\nAlibaba's [Qwen3.5-9B model card](https://huggingface.co/Qwen/Qwen3.5-9B?ref=runtimewire) publishes company-run benchmark comparisons covering knowledge, coding, reasoning, long-context and agent tasks. It also lists a native 262,144-token context window, multimodal inputs and an Apache 2.0 license. Those specifications describe an unusually broad model for its size, though they do not establish the performance of the browser configuration on consumer hardware.\n\nThe distinction matters because the interface packages optimized versions of the smaller models for browser inference. Browser runtime, numerical precision, available memory and GPU implementation can all affect speed and output. A live demo proves that the software stack works. It does not turn Alibaba's broad frontier comparison into an independently reproduced result.\n\nThat leaves a useful, narrower claim: Qwen3.5 Small can provide local multimodal chat through a browser interface, with model choices small enough to target consumer hardware. Developers can test that proposition directly, although the demo supplies no hardware matrix or tokens-per-second results.\n\n### Independent tooling closes the deployment gap\n\nThe demo shows how much of the local-AI contest belongs to independent tooling builders. Labs can publish model cards and weight files, while adoption depends on people who convert those files, support new architectures, optimize kernels and make the result usable outside a research environment.\n\nAlibaba supplied the models. Webml-community maintains the browser deployment, and Lochner brought it to developers' attention. Five and a half months after Qwen3.5 Small shipped, the technical reason to revisit the family is the shorter path from downloadable weights to a local application.", "url": "https://wpnews.pro/news/webml-community-runs-three-qwen3-5-models-locally-in-browsers-with-webgpu", "canonical_source": "https://runtimewire.com/article/alibaba-qwen3-5-small-browser-demo-local-ai", "published_at": "2026-08-18 06:55:06+00:00", "updated_at": "2026-08-18 07:12:36.502131+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-tools", "ai-infrastructure"], "entities": ["webml-community", "Qwen3.5 Small", "Transformers.js", "WebGPU", "Joshua Lochner", "Hugging Face", "Alibaba"], "alternates": {"html": "https://wpnews.pro/news/webml-community-runs-three-qwen3-5-models-locally-in-browsers-with-webgpu", "markdown": "https://wpnews.pro/news/webml-community-runs-three-qwen3-5-models-locally-in-browsers-with-webgpu.md", "text": "https://wpnews.pro/news/webml-community-runs-three-qwen3-5-models-locally-in-browsers-with-webgpu.txt", "jsonld": "https://wpnews.pro/news/webml-community-runs-three-qwen3-5-models-locally-in-browsers-with-webgpu.jsonld"}}