Running LLMs directly in your browser might actually be faster
WebLLM, an open-source library from MLC AI, enables running large language models like Llama 3, Mistral, and Phi-3 directly in the browser via WebGPU, eliminating server-side latency and data privacy …