The idea fits in one sentence: plug a stick into any computer, double-click one
file, and a real language model runs on that machine. No account, no internet,
no install, and nothing left behind when you unplug it.
That sentence took weeks to make true. Not because the AI part is hard, since
llamafile and llama.cpp do the heavy lifting. It took weeks because every
operating system has its own way of quietly refusing to run a program off a
USB stick. Most of those refusals look like something else entirely.
This is the list. Every item happened on real hardware, and the fix for each
is in the kit.
On Windows, llama-server.exe started, printed nothing, opened no window, and
exited with code 0. Code 0 means success. It looked exactly like a
corrupted download, so I re-downloaded it. Twice.
The real cause was three missing Microsoft runtime DLLs: MSVCP140.dll,
VCRUNTIME140.dll and VCRUNTIME140_1.dll. Any machine that has ever
installed a game or Visual Studio has them. A clean machine doesn't, and the
fails before the program can even report an error.
The fix is to ship the DLLs next to the .exe. I didn't just assume that
works: I uninstalled the redistributable in a test VM, confirmed the registry
key and the System32 copies were gone, and ran it again. It worked. (Those DLLs
also have to be redistributed under Microsoft's actual terms, not copied out of
a System32 folder. That cost its own afternoon.)
llamafile is a wonderful trick: one file that runs on macOS, Linux and
Windows. On Windows its file tools worked fine. But every shell command the
model tried to run returned exit -1, while the same command typed by hand
worked perfectly.
Windows has no fork(). llamafile's portability layer emulates it, and
starting a child process is exactly where that emulation breaks. So the drive
carries a second engine just for Windows, llama.cpp's own
llama-server.exe, and the Windows launchers prefer it.
macOS runs a .command file on double-click, and Windows runs a .bat.
Linux never had an equivalent for .sh scripts. And on Ubuntu 26.04, GNOME's
file manager won't launch a .desktop file or an executable script from
anywhere outside the standard app folders. No setting, no "Allow Launching",
no trust flag brings it back. Both launchers on the drive open in a text
editor.
There's no fix from the drive's side. The kit says so plainly and gives you
two ways around it: a one-line terminal command, or a one-time installer that
puts a launcher where GNOME will run it.
When a Linux desktop mounts a FAT32 stick, it marks only files ending in
.exe, .com or .bat as executable. On a FAT32 copy of the drive, that
makes WINDOWS-start-ai.bat the only runnable file on the stick, and running
the Linux launcher fails with "Permission denied".
exFAT doesn't do this. It also lifts FAT32's 4 GB limit on a single file,
which matters once a model is larger than that. Format the stick as exFAT.
I copied the kit, checked it, and everything was there: right sizes, right
checksums. I ejected, plugged it back in, and found a drive full of 0-byte files that still had every filename correct.
The check had read the files back from the operating system's write cache,
not from the stick. It happened twice in one afternoon. The rule now: eject,
replug, then verify.
The Windows launcher opened the chat page with start "" http://.... On a
machine with no default browser, that call never returns, and the whole
launcher freezes with no error.
The fix is to hand the browser launch off through PowerShell, and to always
print the address so you can open it yourself.
My test stick reads at 34 MB/s. The whole model is read off the stick
every time it starts, so a 2.6 GB model takes over a minute before the first
answer. A good USB 3 drive does it in seconds.
On my home server, a bigger "mixture of experts" model was actually faster,
because memory was the bottleneck there. On a USB stick, file size is the only
thing that matters: an 18 GB model would need about nine minutes just to load.
So the kit ships small dense models, and the advice is to upgrade the drive
before you upgrade the model.
Qwen3 is a reasoning model, so it thinks before every answer. Asked to open
the Recycle Bin, it spent 441 tokens deliberating and still got it wrong.
With reasoning off, the same kind of question took about a dozen tokens.
On a small model with a limited context window, that's the difference between
an assistant and something that fills its memory before it answers. Every
launcher turns reasoning off. A related default: llama-server splits its
context between four parallel users. One person on a stick needs one, so the
launchers set it to one and get four times the room.
On macOS, llamafile sometimes crashes inside fork() when the model runs a
shell command. That's an upstream bug I can't fix, so the Mac launcher
restarts the server automatically. I tested it by killing the server myself,
and it came back on the same port within two seconds.
The bug has a second form. Sometimes the forked copy doesn't crash but
deadlocks, spinning at 100% CPU and ignoring the normal stop signal. Two of
them ran for 8.5 hours, held the port, and pushed the next launch onto a
different port without saying so. The launcher now finds these by their
fingerprint (the right port, no parent process, exactly one thread; a live
server has about 17) and cleans them up.
------- divider√¢?".
Version 1.1 adds hearing: a whisper.cpp build (whisperfile) that turns voice
memos and meetings into text, and a push-to-talk page that runs the whole
loop locally. Its decoder fails with failed to read pcm frames: At end on
some files and not others, with no pattern I could see.
The pattern, after a sweep of 40 generated files: it fails whenever the
sample count is an exact multiple of 2048. A browser's audio capture
buffer is 4096 samples, so every recording from the voice page hit it. It
also fails on WAVs that carry a metadata chunk, which ffmpeg adds when the
source is an .m4a. The verb now re-encodes everything, strips metadata, and
pads one extra sample if the count lands on the wrong number. On a Mac with
no ffmpeg it falls back to afconvert and rewrites the 68-byte header it
produces into the plain 44-byte one the decoder expects.
Two more from the same release: spd-say -w on a headless Linux box blocks
forever, so the talk-back verb now gives it a time budget; and a server
started from a tool that runs at nice 5 gets its child processes parked on
efficiency cores under load, which made a 3-second transcription take 37.
Seven models with a picker at launch, from Qwen3-1.7B (fast, runs on
anything) to Qwen3-8B (smartest, wants about 16 GB of RAM), plus a Vision model that reads the photos, receipts and screenshots you attach. It
transcribes audio and reads answers aloud, and a voice page lets you hold a
button, ask, and hear the reply, with the microphone audio going to a server
on localhost and nowhere else. Launchers for macOS, Windows and Linux, a
files-only mode, and an agent mode that asks your permission in the browser
before every shell command it runs. A memory file lives on the drive, so
the assistant remembers you from machine to machine.
It was tested on an Apple Silicon Mac, Ubuntu 26.04, and Windows Server 2025
running off the physical stick. The models are small: they're useful for
writing, summarising, editing files, reading a receipt and running a computer
from plain English, but they're not frontier models and they'll sometimes be confidently wrong. Nothing on a USB stick is a frontier model.
If you'd rather build it yourself, the build kit downloads every model and engine straight from the people who publish them and checks each against the
publisher's SHA-256. If you just want it working, there's a ready-made 13 GB
drive image you copy onto a stick.
Portable AI on a Drive, $59: https://primeagent2.gumroad.com/l/objkjr Launch code USB44 takes $15 off until 27 October.
Originally published on Gumroad.