[ Video ] How To Set Up Local AI on a 16GB MacBook A developer published a video and written guide showing how to run local AI models on a 16GB M1 MacBook Pro using the mlx-serve command-line utility, which is optimized for Apple's Metal architecture. The guide covers three ways to consume served models: the Grok Build CLI, the OpenWeb UI chat wrapper, and the notarized MLX-Serve desktop app, which also supports image, voice, music and video generation plus agent definitions and recurring tasks. One of my YouTube subscribers asked, in a comment, how to set up a 16GB M1 MacBook Pro so it runs models in a more comfortable way, not necessarily in the terminal. This came as a reaction to my local models demo videos, which seem to be picking up lately. People enjoy that content, but they want something more visual to play with them. So, I made a new video, specifically for this. Just watch it above, if you’re the visual type. If not, read on for a detailed breakdown on how it’s done. The basic understanding of running local AI is that you first have to “serve” the model, and that is done usually in the terminal. After this step, you can start “consuming” the model, via a web interface, or even a desktop app. Serving the model is done, on a mac, with an utility called mlx-serve . It’s a solid command line utility, with a lot of arguments, which sometimes may be overwhelming. But the TLDR here is that you only need to remember 2 commands: mlx-serve run mlx-serve serve Depending on whether or not a model is already local, it also downloads the model on your machine. We are not using ollama or other more established binaries because mlx is specifically optimized for the Metal architecture of your Mac. Simply put, the models will run faster and with less resources. Serving the model is already 50% of the job. From there on, I explore 3 ways to consume this. Grok Build is my default CLI these days – I’m vintage, I know, and I spend 90% of my time in the terminal. I already talked about how to set up free models both local and remote, on OpenRouter on this post https://dragosroua.com/how-i-use-grok-build-to-run-free-models-on-top-of-grok-4-6/ . The video is just a very short version of that. The TLDR is that you will edit Grok’s config.toml by adding the model characteristics, and then you’re just picking the model mid-session with /models inside Grok Build. It’s not complicated, if you like staying the terminal, and, like I said, that’s my goto approach when using AI in general. OpenWeb UI is a chat wrapper on top of the local model, but it offers a relatively rich experience. It’s very simple to install, and it will offers you a familiar, ChatGPT style interface. You can switch models or adjust the model parameters, and generally do whatever you do with a standard ChatGPT interface, only it uses a local model. Use it if you’re spending most of the time in the browser. The most complex and rich experience comes from MLX-Serve, which is a notarized by Apple app, built on top of the command line utility mentioned at the beginning of this article. It is by far the nicest, most elegant way to run local AI on your machine and it has all the pre-requisites to become a more established app for day to day AI tasks. On top of the chat experience, you also get to generate images, voice text to speech , music and even videos. There’s solid support for defining agents too, and you get recurring tasks as well. That’s my recommendation for anyone who want to try local AI on a Mac. All the links for the above programs are in the video, feel free to check it out. As usual, if you like this kind of content, like, share and subscribe to my YouTube channel, you will help local AI grow.