# I put seven LLMs on a USB stick. Here is everything that broke.

> Source: <https://dev.to/alfred_odong_322108a5cc3d/i-put-seven-llms-on-a-usb-stick-here-is-everything-that-broke-39cc>
> Published: 2026-10-08 06:09:51+00:00

The idea fits in one sentence: plug a stick into any computer, double-click one

file, and a real language model runs on that machine. No account, no internet,

no install, and nothing left behind when you unplug it.

That sentence took weeks to make true. Not because the AI part is hard, since

llamafile and llama.cpp do the heavy lifting. It took weeks because every

operating system has its own way of quietly refusing to run a program off a

USB stick. Most of those refusals look like something else entirely.

This is the list. Every item happened on real hardware, and the fix for each

is in the kit.

On Windows, `llama-server.exe` started, printed nothing, opened no window, and

exited with **code 0**. Code 0 means success. It looked exactly like a

corrupted download, so I re-downloaded it. Twice.

The real cause was three missing Microsoft runtime DLLs: `MSVCP140.dll`,

`VCRUNTIME140.dll` and `VCRUNTIME140_1.dll`. Any machine that has ever

installed a game or Visual Studio has them. A clean machine doesn't, and the

loader fails before the program can even report an error.

The fix is to ship the DLLs next to the `.exe`. I didn't just assume that

works: I uninstalled the redistributable in a test VM, confirmed the registry

key and the System32 copies were gone, and ran it again. It worked. (Those DLLs

also have to be redistributed under Microsoft's actual terms, not copied out of

a System32 folder. That cost its own afternoon.)

llamafile is a wonderful trick: one file that runs on macOS, Linux and

Windows. On Windows its file tools worked fine. But every shell command the

model tried to run returned `exit -1`, while the same command typed by hand

worked perfectly.

Windows has no `fork()`. llamafile's portability layer emulates it, and

starting a child process is exactly where that emulation breaks. So the drive

carries a second engine just for Windows, llama.cpp's own

`llama-server.exe`, and the Windows launchers prefer it.

macOS runs a `.command` file on double-click, and Windows runs a `.bat`.

Linux never had an equivalent for `.sh` scripts. And on Ubuntu 26.04, GNOME's

file manager won't launch a `.desktop` file *or* an executable script from

anywhere outside the standard app folders. No setting, no "Allow Launching",

no trust flag brings it back. Both launchers on the drive open in a text

editor.

There's no fix from the drive's side. The kit says so plainly and gives you

two ways around it: a one-line terminal command, or a one-time installer that

puts a launcher where GNOME will run it.

When a Linux desktop mounts a FAT32 stick, it marks only files ending in

`.exe`, `.com` or `.bat` as executable. On a FAT32 copy of the drive, that

makes `WINDOWS-start-ai.bat` the only runnable file on the stick, and running

the Linux launcher fails with "Permission denied".

exFAT doesn't do this. It also lifts FAT32's 4 GB limit on a single file,

which matters once a model is larger than that. Format the stick as exFAT.

I copied the kit, checked it, and everything was there: right sizes, right

checksums. I ejected, plugged it back in, and found a drive full of **0-byte files** that still had every filename correct.

The check had read the files back from the operating system's write cache,

not from the stick. It happened twice in one afternoon. The rule now: eject,

replug, *then* verify.

The Windows launcher opened the chat page with `start "" http://...`. On a

machine with no default browser, that call never returns, and the whole

launcher freezes with no error.

The fix is to hand the browser launch off through PowerShell, and to always

print the address so you can open it yourself.

My test stick reads at **34 MB/s**. The whole model is read off the stick

every time it starts, so a 2.6 GB model takes over a minute before the first

answer. A good USB 3 drive does it in seconds.

On my home server, a bigger "mixture of experts" model was actually *faster*,

because memory was the bottleneck there. On a USB stick, file size is the only

thing that matters: an 18 GB model would need about nine minutes just to load.

So the kit ships small dense models, and the advice is to upgrade the drive

before you upgrade the model.

Qwen3 is a reasoning model, so it thinks before every answer. Asked to open

the Recycle Bin, it spent **441 tokens** deliberating and still got it wrong.

With reasoning off, the same kind of question took about a dozen tokens.

On a small model with a limited context window, that's the difference between

an assistant and something that fills its memory before it answers. Every

launcher turns reasoning off. A related default: llama-server splits its

context between four parallel users. One person on a stick needs one, so the

launchers set it to one and get four times the room.

On macOS, llamafile sometimes crashes inside `fork()` when the model runs a

shell command. That's an upstream bug I can't fix, so the Mac launcher

restarts the server automatically. I tested it by killing the server myself,

and it came back on the same port within two seconds.

The bug has a second form. Sometimes the forked copy doesn't crash but

deadlocks, spinning at 100% CPU and ignoring the normal stop signal. Two of

them ran for 8.5 hours, held the port, and pushed the next launch onto a

different port without saying so. The launcher now finds these by their

fingerprint (the right port, no parent process, exactly one thread; a live

server has about 17) and cleans them up.

`-------` divider`√¢?"`.
Version 1.1 adds hearing: a whisper.cpp build (whisperfile) that turns voice

memos and meetings into text, and a push-to-talk page that runs the whole

loop locally. Its decoder fails with `failed to read pcm frames: At end` on

some files and not others, with no pattern I could see.

The pattern, after a sweep of 40 generated files: it fails whenever the

sample count is an exact multiple of **2048**. A browser's audio capture

buffer is 4096 samples, so every recording from the voice page hit it. It

also fails on WAVs that carry a metadata chunk, which ffmpeg adds when the

source is an `.m4a`. The verb now re-encodes everything, strips metadata, and

pads one extra sample if the count lands on the wrong number. On a Mac with

no ffmpeg it falls back to `afconvert` and rewrites the 68-byte header it

produces into the plain 44-byte one the decoder expects.

Two more from the same release: `spd-say -w` on a headless Linux box blocks

forever, so the talk-back verb now gives it a time budget; and a server

started from a tool that runs at `nice 5` gets its child processes parked on

efficiency cores under load, which made a 3-second transcription take 37.

Seven models with a picker at launch, from Qwen3-1.7B (fast, runs on

anything) to Qwen3-8B (smartest, wants about 16 GB of RAM), plus a Vision

model that reads the photos, receipts and screenshots you attach. It

transcribes audio and reads answers aloud, and a voice page lets you hold a

button, ask, and hear the reply, with the microphone audio going to a server

on localhost and nowhere else. Launchers for macOS, Windows and Linux, a

files-only mode, and an agent mode that asks your permission in the browser

before **every** shell command it runs. A memory file lives on the drive, so

the assistant remembers you from machine to machine.

It was tested on an Apple Silicon Mac, Ubuntu 26.04, and Windows Server 2025

running off the physical stick. The models are small: they're useful for

writing, summarising, editing files, reading a receipt and running a computer

from plain English, but they're not frontier models and they'll sometimes be

confidently wrong. Nothing on a USB stick is a frontier model.

If you'd rather build it yourself, the build kit downloads every model and

engine straight from the people who publish them and checks each against the

publisher's SHA-256. If you just want it working, there's a ready-made 13 GB

drive image you copy onto a stick.

**Portable AI on a Drive, $59:** [https://primeagent2.gumroad.com/l/objkjr](https://primeagent2.gumroad.com/l/objkjr)

Launch code **USB44** takes $15 off until 27 October.

*Originally published on [Gumroad](https://primeagent2.gumroad.com/p/i-put-seven-llms-on-a-usb-stick-here-is-everything-that-broke).*
