cd /news/artificial-intelligence/stop-sending-your-vitals-to-the-clou… · home topics artificial-intelligence article
[ARTICLE · art-98388] src=dev.to ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Stop Sending Your Vitals to the Cloud: Running Llama-3 Locally in the Browser with WebLLM & WebGPU 🥑

A developer has created a privacy-first health log analyzer that runs Llama-3-8B locally in the browser using WebLLM and WebGPU, ensuring no personal health information leaves the device. The system combines Transformers.js for entity extraction and WebLLM for reasoning, with data stored in IndexedDB. The project demonstrates a shift from client-server to edge AI for sensitive applications.

read4 min views1 publishedAug 16, 2026

Privacy is the ultimate "final boss" in HealthTech. When users record sensitive medical logs, the last thing they want is their data being used to train a massive corporate model. Today, we are pushing the boundaries of Edge AI by building a 100% private, client-side health log analyzer. By leveraging WebGPU acceleration and WebLLM, we can run a full Llama-3 instance directly in the browser.

In this tutorial, we will explore how to combine Transformers.js for lightweight feature extraction and WebLLM for complex reasoning. This approach ensures that your privacy-first health apps remain performant without a single byte of personal health information (PHI) ever leaving the user's device. Let’s dive into the world of local LLM inference and browser-based machine learning! 🚀

Traditional AI apps follow a Client-Server model. We are flipping the script. Our architecture keeps the data, the model, and the compute inside the browser's sandbox.

graph TD
    A[User Inputs Health Log] --> B{Local Processing}
    B --> C[Transformers.js: Entity Extraction]
    B --> D[WebLLM: Llama-3-8B Reasoning]
    C --> E[Structured Health Data]
    D --> F[Clinical Insights & Summary]
    E --> G[IndexedDB: Local Storage]
    F --> G
    G --> H[Privacy-Safe UI View]
    style B fill:#f9f,stroke:#333,stroke-width:4px

Before we start coding, ensure your environment meets these requirements:

WebLLM is a high-performance in-browser LLM inference engine. It uses the WebGPU API to execute model weights compiled with TVM.

First, install the dependency:

npm install @mlc-ai/web-llm

Now, let's create a hook to manage our Llama-3 instance. We’ll use the Llama-3-8B-Instruct-q4f16_1-MLC

variant, which is optimized for 4-bit quantization to fit in browser memory.

import { useState, useEffect } from 'react';
import * as webllm from "@mlc-ai/web-llm";

export function useWebLLM() {
  const [engine, setEngine] = useState<webllm.MLCEngine | null>(null);
  const [Progress, setProgress] = useState(0);

  const initEngine = async () => {
    const engine = new webllm.MLCEngine();

    // Callback to track model down/ progress
    engine.setInitProgressCallback((report) => {
      setProgress(Math.round(report.progress * 100));
      console.log(report.text);
    });

    const selectedModel = "Llama-3-8B-Instruct-q4f16_1-MLC";
    await engine.reload(selectedModel);
    setEngine(engine);
  };

  return { engine, initEngine, Progress };
}

While Llama-3 handles the heavy reasoning, we can use Transformers.js for fast, local Named Entity Recognition (NER). This is great for identifying medications or symptoms before passing them to the LLM.

import { pipeline } from '@xenova/transformers';

const analyzeLogBasics = async (text) => {
  // Use a tiny, efficient model for fast extraction
  const extractor = await pipeline('token-classification', 'Xenova/bert-base-NER');
  const results = await extractor(text);

  // Filter for medical-related entities locally
  return results.filter(entity => ['MED', 'SYMPTOM'].includes(entity.entity));
};

When building production-grade healthcare applications, simply running a model isn't enough. You need to handle state management, local encryption, and sophisticated prompt engineering.

For a deeper dive into production-ready Edge AI patterns and advanced security protocols for health data, I highly recommend checking out the technical deep-dives at WellAlly Blog. They offer incredible resources on how to bridge the gap between "cool browser demos" and "HIPAA-compliant local software."

Now, let's combine everything into a React component. The user types their log, we extract entities, and then Llama-3 provides a clinical summary—all on the GPU.

import React, { useState } from 'react';
import { useWebLLM } from './hooks/useWebLLM';

const HealthAnalyzer = () => {
  const { engine, initEngine, Progress } = useWebLLM();
  const [input, setInput] = useState("");
  const [output, setOutput] = useState("");

  const handleAnalyze = async () => {
    if (!engine) return;

    const messages = [
      { role: "system", content: "You are a private health assistant. Analyze the user's log for potential trends. Keep it professional." },
      { role: "user", content: input }
    ];

    const reply = await engine.chat.completions.create({ messages });
    setOutput(reply.choices[0].message.content);
  };

  return (
    <div className="p-8 max-w-2xl mx-auto">
      <h2 className="text-2xl font-bold mb-4">Local Health Log 🩺</h2>

      {!engine ? (
        <button 
          onClick={initEngine}
          className="bg-blue-600 text-white px-4 py-2 rounded"
        >
          Load Llama-3 ({Progress}%)
        </button>
      ) : (
        <div className="space-y-4">
          <textarea 
            className="w-full border p-2"
            placeholder="e.g., Feeling dizzy after taking 20mg Lisinopril..."
            onChange={(e) => setInput(e.target.value)}
          />
          <button 
            onClick={handleAnalyze}
            className="bg-green-600 text-white px-4 py-2 rounded"
          >
            Analyze Privately
          </button>
          <div className="mt-4 p-4 bg-gray-100 rounded">
            <strong>Insight:</strong> {output}
          </div>
        </div>
      )}
    </div>
  );
};

Running Llama-3 in the browser isn't just a party trick; it's a paradigm shift for Edge AI and Privacy. By using WebLLM and WebGPU, we give power back to the users while maintaining the "magic" of LLMs.

Are you ready to move your AI workloads to the edge? Let me know in the comments if you've tried running local models! And don't forget to visit wellally.tech/blog for more advanced AI architecture guides. 🥑💻

Happy coding!

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @webllm 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/stop-sending-your-vi…] indexed:0 read:4min 2026-08-16 ·