cd /news/large-language-models/i-put-seven-llms-on-a-usb-stick-here… · home › topics › large-language-models › article
[ARTICLE · art-147383] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=· neutral

I put seven LLMs on a USB stick. Here is everything that broke.

A developer documented building a portable USB stick that runs seven local language models offline via llamafile and llama.cpp, cataloguing the OS-level failures encountered along the way. The writeup details Windows missing MSVCP140.dll and VCRUNTIME140.dll runtime errors that exit with code 0, llamafile's broken fork() emulation for shell commands on Windows, GNOME's refusal to launch .desktop or script files on Ubuntu 26.04, FAT32 executable-bit restrictions, and 0-byte file corruption caused by verifying against the OS write cache instead of the stick. The kit's fixes include shipping the runtime DLLs beside the executable, preferring llama-server.exe on Windows, formatting as exFAT, and ejecting and replugging before verification.

by read7 min views1 publishedOct 8, 2026

The idea fits in one sentence: plug a stick into any computer, double-click one

file, and a real language model runs on that machine. No account, no internet,

no install, and nothing left behind when you unplug it.

That sentence took weeks to make true. Not because the AI part is hard, since

llamafile and llama.cpp do the heavy lifting. It took weeks because every

operating system has its own way of quietly refusing to run a program off a

USB stick. Most of those refusals look like something else entirely.

This is the list. Every item happened on real hardware, and the fix for each

is in the kit.

On Windows, llama-server.exe started, printed nothing, opened no window, and

exited with code 0. Code 0 means success. It looked exactly like a

corrupted download, so I re-downloaded it. Twice.

The real cause was three missing Microsoft runtime DLLs: MSVCP140.dll,

VCRUNTIME140.dll and VCRUNTIME140_1.dll. Any machine that has ever

installed a game or Visual Studio has them. A clean machine doesn't, and the

fails before the program can even report an error.

The fix is to ship the DLLs next to the .exe. I didn't just assume that

works: I uninstalled the redistributable in a test VM, confirmed the registry

key and the System32 copies were gone, and ran it again. It worked. (Those DLLs

also have to be redistributed under Microsoft's actual terms, not copied out of

a System32 folder. That cost its own afternoon.)

llamafile is a wonderful trick: one file that runs on macOS, Linux and

Windows. On Windows its file tools worked fine. But every shell command the

model tried to run returned exit -1, while the same command typed by hand

worked perfectly.

Windows has no fork(). llamafile's portability layer emulates it, and

starting a child process is exactly where that emulation breaks. So the drive

carries a second engine just for Windows, llama.cpp's own

llama-server.exe, and the Windows launchers prefer it.

macOS runs a .command file on double-click, and Windows runs a .bat.

Linux never had an equivalent for .sh scripts. And on Ubuntu 26.04, GNOME's

file manager won't launch a .desktop file or an executable script from

anywhere outside the standard app folders. No setting, no "Allow Launching",

no trust flag brings it back. Both launchers on the drive open in a text

editor.

There's no fix from the drive's side. The kit says so plainly and gives you

two ways around it: a one-line terminal command, or a one-time installer that

puts a launcher where GNOME will run it.

When a Linux desktop mounts a FAT32 stick, it marks only files ending in

.exe, .com or .bat as executable. On a FAT32 copy of the drive, that

makes WINDOWS-start-ai.bat the only runnable file on the stick, and running

the Linux launcher fails with "Permission denied".

exFAT doesn't do this. It also lifts FAT32's 4 GB limit on a single file,

which matters once a model is larger than that. Format the stick as exFAT.

I copied the kit, checked it, and everything was there: right sizes, right

checksums. I ejected, plugged it back in, and found a drive full of 0-byte files that still had every filename correct.

The check had read the files back from the operating system's write cache,

not from the stick. It happened twice in one afternoon. The rule now: eject,

replug, then verify.

The Windows launcher opened the chat page with start "" http://.... On a

machine with no default browser, that call never returns, and the whole

launcher freezes with no error.

The fix is to hand the browser launch off through PowerShell, and to always

print the address so you can open it yourself.

My test stick reads at 34 MB/s. The whole model is read off the stick

every time it starts, so a 2.6 GB model takes over a minute before the first

answer. A good USB 3 drive does it in seconds.

On my home server, a bigger "mixture of experts" model was actually faster,

because memory was the bottleneck there. On a USB stick, file size is the only

thing that matters: an 18 GB model would need about nine minutes just to load.

So the kit ships small dense models, and the advice is to upgrade the drive

before you upgrade the model.

Qwen3 is a reasoning model, so it thinks before every answer. Asked to open

the Recycle Bin, it spent 441 tokens deliberating and still got it wrong.

With reasoning off, the same kind of question took about a dozen tokens.

On a small model with a limited context window, that's the difference between

an assistant and something that fills its memory before it answers. Every

launcher turns reasoning off. A related default: llama-server splits its

context between four parallel users. One person on a stick needs one, so the

launchers set it to one and get four times the room.

On macOS, llamafile sometimes crashes inside fork() when the model runs a

shell command. That's an upstream bug I can't fix, so the Mac launcher

restarts the server automatically. I tested it by killing the server myself,

and it came back on the same port within two seconds.

The bug has a second form. Sometimes the forked copy doesn't crash but

deadlocks, spinning at 100% CPU and ignoring the normal stop signal. Two of

them ran for 8.5 hours, held the port, and pushed the next launch onto a

different port without saying so. The launcher now finds these by their

fingerprint (the right port, no parent process, exactly one thread; a live

server has about 17) and cleans them up.

------- divider√¢?". Version 1.1 adds hearing: a whisper.cpp build (whisperfile) that turns voice

memos and meetings into text, and a push-to-talk page that runs the whole

loop locally. Its decoder fails with failed to read pcm frames: At end on

some files and not others, with no pattern I could see.

The pattern, after a sweep of 40 generated files: it fails whenever the

sample count is an exact multiple of 2048. A browser's audio capture

buffer is 4096 samples, so every recording from the voice page hit it. It

also fails on WAVs that carry a metadata chunk, which ffmpeg adds when the

source is an .m4a. The verb now re-encodes everything, strips metadata, and

pads one extra sample if the count lands on the wrong number. On a Mac with

no ffmpeg it falls back to afconvert and rewrites the 68-byte header it

produces into the plain 44-byte one the decoder expects.

Two more from the same release: spd-say -w on a headless Linux box blocks

forever, so the talk-back verb now gives it a time budget; and a server

started from a tool that runs at nice 5 gets its child processes parked on

efficiency cores under load, which made a 3-second transcription take 37.

Seven models with a picker at launch, from Qwen3-1.7B (fast, runs on

anything) to Qwen3-8B (smartest, wants about 16 GB of RAM), plus a Vision model that reads the photos, receipts and screenshots you attach. It

transcribes audio and reads answers aloud, and a voice page lets you hold a

button, ask, and hear the reply, with the microphone audio going to a server

on localhost and nowhere else. Launchers for macOS, Windows and Linux, a

files-only mode, and an agent mode that asks your permission in the browser

before every shell command it runs. A memory file lives on the drive, so

the assistant remembers you from machine to machine.

It was tested on an Apple Silicon Mac, Ubuntu 26.04, and Windows Server 2025

running off the physical stick. The models are small: they're useful for

writing, summarising, editing files, reading a receipt and running a computer

from plain English, but they're not frontier models and they'll sometimes be confidently wrong. Nothing on a USB stick is a frontier model.

If you'd rather build it yourself, the build kit downloads every model and engine straight from the people who publish them and checks each against the

publisher's SHA-256. If you just want it working, there's a ready-made 13 GB

drive image you copy onto a stick.

Portable AI on a Drive, $59: https://primeagent2.gumroad.com/l/objkjr Launch code USB44 takes $15 off until 27 October.

Originally published on Gumroad.

── more in #large-language-models 4 stories · sorted by recency
── more on @llamafile 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/i-put-seven-llms-on-…] indexed:0 read:7min 2026-10-08 · —