Can we run large AI models on computers that don’t have enough RAM to hold their weights?
I’m building Cortex, an experimental open-source AI inference system focused on streaming model weights from storage rather than requiring the entire model to remain resident in memory.
The long-term goal is to make larger AI models usable on older, inexpensive, CPU-only computers.
Working with BitNet, GGUF, and llama.cpp-based experiments, I’ve made progress on:
CPU-only inference on an AMD Athlon Silver 3050U.
A baseline of approximately 10 tokens/second.
External tiled output projection with verified matching logits.
Streaming input and output tensor data.
Reusable 16 MiB tile buffering.
Bit-for-bit verification of selected streaming computations against baseline results.
These are experimental milestones, not yet a complete low-memory inference solution.
The central challenge is separating model storage from model residency.
Even when weights are read externally during computation, parts of the existing inference runtime still expect tensors to be registered and resident in memory.
I want to solve this properly, not simply disguise memory usage with swap or zram.
llama.cpp / GGML developers: Help redesign tensor and execution to support genuinely nonresident weights.
Systems programmers: Help optimize asynchronous I/O, prefetching, buffering, and memory residency.
Quantization specialists: Help support BitNet and other quantized model architectures.
Performance engineers: Help benchmark the tradeoffs between RAM, storage bandwidth, CPU utilization, and tokens per second.
AI researchers: Explore whether this architecture could generalize across different transformer models.
I want Cortex to evolve into a model-independent streaming inference runtime capable of running different AI architectures with limited memory.
The goal isn’t to claim that streaming magically eliminates computation or storage bottlenecks. It’s to discover how far careful execution scheduling and memory management can take us.
I’m looking for collaborators who enjoy difficult systems problems and want to help turn this experimental work into a reproducible, usable open-source project.
GitHub: [Insert Cortex repository URL]
If you’ve worked on llama.cpp, GGML, BitNet, out-of-core computation, model off, or inference runtimes, I’d really appreciate your input.
Let’s see how far we can push AI inference on hardware that most people have already written off.
My github profile is @ jddchatham · GitHub