{"slug": "i-put-seven-llms-on-a-usb-stick-here-is-everything-that-broke", "title": "I put seven LLMs on a USB stick. Here is everything that broke.", "summary": "A developer documented building a portable USB stick that runs seven local language models offline via llamafile and llama.cpp, cataloguing the OS-level failures encountered along the way. The writeup details Windows missing MSVCP140.dll and VCRUNTIME140.dll runtime errors that exit with code 0, llamafile's broken fork() emulation for shell commands on Windows, GNOME's refusal to launch .desktop or script files on Ubuntu 26.04, FAT32 executable-bit restrictions, and 0-byte file corruption caused by verifying against the OS write cache instead of the stick. The kit's fixes include shipping the runtime DLLs beside the executable, preferring llama-server.exe on Windows, formatting as exFAT, and ejecting and replugging before verification.", "body_md": "The idea fits in one sentence: plug a stick into any computer, double-click one\n\nfile, and a real language model runs on that machine. No account, no internet,\n\nno install, and nothing left behind when you unplug it.\n\nThat sentence took weeks to make true. Not because the AI part is hard, since\n\nllamafile and llama.cpp do the heavy lifting. It took weeks because every\n\noperating system has its own way of quietly refusing to run a program off a\n\nUSB stick. Most of those refusals look like something else entirely.\n\nThis is the list. Every item happened on real hardware, and the fix for each\n\nis in the kit.\n\nOn Windows, `llama-server.exe` started, printed nothing, opened no window, and\n\nexited with **code 0**. Code 0 means success. It looked exactly like a\n\ncorrupted download, so I re-downloaded it. Twice.\n\nThe real cause was three missing Microsoft runtime DLLs: `MSVCP140.dll`,\n\n`VCRUNTIME140.dll` and `VCRUNTIME140_1.dll`. Any machine that has ever\n\ninstalled a game or Visual Studio has them. A clean machine doesn't, and the\n\nloader fails before the program can even report an error.\n\nThe fix is to ship the DLLs next to the `.exe`. I didn't just assume that\n\nworks: I uninstalled the redistributable in a test VM, confirmed the registry\n\nkey and the System32 copies were gone, and ran it again. It worked. (Those DLLs\n\nalso have to be redistributed under Microsoft's actual terms, not copied out of\n\na System32 folder. That cost its own afternoon.)\n\nllamafile is a wonderful trick: one file that runs on macOS, Linux and\n\nWindows. On Windows its file tools worked fine. But every shell command the\n\nmodel tried to run returned `exit -1`, while the same command typed by hand\n\nworked perfectly.\n\nWindows has no `fork()`. llamafile's portability layer emulates it, and\n\nstarting a child process is exactly where that emulation breaks. So the drive\n\ncarries a second engine just for Windows, llama.cpp's own\n\n`llama-server.exe`, and the Windows launchers prefer it.\n\nmacOS runs a `.command` file on double-click, and Windows runs a `.bat`.\n\nLinux never had an equivalent for `.sh` scripts. And on Ubuntu 26.04, GNOME's\n\nfile manager won't launch a `.desktop` file *or* an executable script from\n\nanywhere outside the standard app folders. No setting, no \"Allow Launching\",\n\nno trust flag brings it back. Both launchers on the drive open in a text\n\neditor.\n\nThere's no fix from the drive's side. The kit says so plainly and gives you\n\ntwo ways around it: a one-line terminal command, or a one-time installer that\n\nputs a launcher where GNOME will run it.\n\nWhen a Linux desktop mounts a FAT32 stick, it marks only files ending in\n\n`.exe`, `.com` or `.bat` as executable. On a FAT32 copy of the drive, that\n\nmakes `WINDOWS-start-ai.bat` the only runnable file on the stick, and running\n\nthe Linux launcher fails with \"Permission denied\".\n\nexFAT doesn't do this. It also lifts FAT32's 4 GB limit on a single file,\n\nwhich matters once a model is larger than that. Format the stick as exFAT.\n\nI copied the kit, checked it, and everything was there: right sizes, right\n\nchecksums. I ejected, plugged it back in, and found a drive full of **0-byte files** that still had every filename correct.\n\nThe check had read the files back from the operating system's write cache,\n\nnot from the stick. It happened twice in one afternoon. The rule now: eject,\n\nreplug, *then* verify.\n\nThe Windows launcher opened the chat page with `start \"\" http://...`. On a\n\nmachine with no default browser, that call never returns, and the whole\n\nlauncher freezes with no error.\n\nThe fix is to hand the browser launch off through PowerShell, and to always\n\nprint the address so you can open it yourself.\n\nMy test stick reads at **34 MB/s**. The whole model is read off the stick\n\nevery time it starts, so a 2.6 GB model takes over a minute before the first\n\nanswer. A good USB 3 drive does it in seconds.\n\nOn my home server, a bigger \"mixture of experts\" model was actually *faster*,\n\nbecause memory was the bottleneck there. On a USB stick, file size is the only\n\nthing that matters: an 18 GB model would need about nine minutes just to load.\n\nSo the kit ships small dense models, and the advice is to upgrade the drive\n\nbefore you upgrade the model.\n\nQwen3 is a reasoning model, so it thinks before every answer. Asked to open\n\nthe Recycle Bin, it spent **441 tokens** deliberating and still got it wrong.\n\nWith reasoning off, the same kind of question took about a dozen tokens.\n\nOn a small model with a limited context window, that's the difference between\n\nan assistant and something that fills its memory before it answers. Every\n\nlauncher turns reasoning off. A related default: llama-server splits its\n\ncontext between four parallel users. One person on a stick needs one, so the\n\nlaunchers set it to one and get four times the room.\n\nOn macOS, llamafile sometimes crashes inside `fork()` when the model runs a\n\nshell command. That's an upstream bug I can't fix, so the Mac launcher\n\nrestarts the server automatically. I tested it by killing the server myself,\n\nand it came back on the same port within two seconds.\n\nThe bug has a second form. Sometimes the forked copy doesn't crash but\n\ndeadlocks, spinning at 100% CPU and ignoring the normal stop signal. Two of\n\nthem ran for 8.5 hours, held the port, and pushed the next launch onto a\n\ndifferent port without saying so. The launcher now finds these by their\n\nfingerprint (the right port, no parent process, exactly one thread; a live\n\nserver has about 17) and cleans them up.\n\n`-------` divider`√¢?\"`.\nVersion 1.1 adds hearing: a whisper.cpp build (whisperfile) that turns voice\n\nmemos and meetings into text, and a push-to-talk page that runs the whole\n\nloop locally. Its decoder fails with `failed to read pcm frames: At end` on\n\nsome files and not others, with no pattern I could see.\n\nThe pattern, after a sweep of 40 generated files: it fails whenever the\n\nsample count is an exact multiple of **2048**. A browser's audio capture\n\nbuffer is 4096 samples, so every recording from the voice page hit it. It\n\nalso fails on WAVs that carry a metadata chunk, which ffmpeg adds when the\n\nsource is an `.m4a`. The verb now re-encodes everything, strips metadata, and\n\npads one extra sample if the count lands on the wrong number. On a Mac with\n\nno ffmpeg it falls back to `afconvert` and rewrites the 68-byte header it\n\nproduces into the plain 44-byte one the decoder expects.\n\nTwo more from the same release: `spd-say -w` on a headless Linux box blocks\n\nforever, so the talk-back verb now gives it a time budget; and a server\n\nstarted from a tool that runs at `nice 5` gets its child processes parked on\n\nefficiency cores under load, which made a 3-second transcription take 37.\n\nSeven models with a picker at launch, from Qwen3-1.7B (fast, runs on\n\nanything) to Qwen3-8B (smartest, wants about 16 GB of RAM), plus a Vision\n\nmodel that reads the photos, receipts and screenshots you attach. It\n\ntranscribes audio and reads answers aloud, and a voice page lets you hold a\n\nbutton, ask, and hear the reply, with the microphone audio going to a server\n\non localhost and nowhere else. Launchers for macOS, Windows and Linux, a\n\nfiles-only mode, and an agent mode that asks your permission in the browser\n\nbefore **every** shell command it runs. A memory file lives on the drive, so\n\nthe assistant remembers you from machine to machine.\n\nIt was tested on an Apple Silicon Mac, Ubuntu 26.04, and Windows Server 2025\n\nrunning off the physical stick. The models are small: they're useful for\n\nwriting, summarising, editing files, reading a receipt and running a computer\n\nfrom plain English, but they're not frontier models and they'll sometimes be\n\nconfidently wrong. Nothing on a USB stick is a frontier model.\n\nIf you'd rather build it yourself, the build kit downloads every model and\n\nengine straight from the people who publish them and checks each against the\n\npublisher's SHA-256. If you just want it working, there's a ready-made 13 GB\n\ndrive image you copy onto a stick.\n\n**Portable AI on a Drive, $59:** [https://primeagent2.gumroad.com/l/objkjr](https://primeagent2.gumroad.com/l/objkjr)\n\nLaunch code **USB44** takes $15 off until 27 October.\n\n*Originally published on [Gumroad](https://primeagent2.gumroad.com/p/i-put-seven-llms-on-a-usb-stick-here-is-everything-that-broke).*", "url": "https://wpnews.pro/news/i-put-seven-llms-on-a-usb-stick-here-is-everything-that-broke", "canonical_source": "https://dev.to/alfred_odong_322108a5cc3d/i-put-seven-llms-on-a-usb-stick-here-is-everything-that-broke-39cc", "published_at": "2026-10-08 06:09:51+00:00", "updated_at": "2026-10-08 06:17:29.491177+00:00", "lang": "en", "topics": ["large-language-models", "ai-tools", "ai-infrastructure", "developer-tools"], "entities": ["llamafile", "llama.cpp", "Microsoft", "Windows", "macOS", "Linux", "GNOME", "Ubuntu"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/i-put-seven-llms-on-a-usb-stick-here-is-everything-that-broke", "markdown": "https://wpnews.pro/news/i-put-seven-llms-on-a-usb-stick-here-is-everything-that-broke.md", "text": "https://wpnews.pro/news/i-put-seven-llms-on-a-usb-stick-here-is-everything-that-broke.txt", "jsonld": "https://wpnews.pro/news/i-put-seven-llms-on-a-usb-stick-here-is-everything-that-broke.jsonld"}}