{"slug": "range-open-a-1-tb-ai-model-in-3-seconds-without-downloading-it", "title": "Range – open a 1 TB AI model in 3 seconds without downloading it", "summary": "Software engineer Andrey Grehov released Range, an open-source tool that opens a shell in a container image, Hugging Face repository, or S3/HTTP environment by reading only the bytes a program touches rather than downloading the whole source. In benchmarks run on EC2 m6i.large in us-east-1 on 28 September 2026, Range answered a llama.cpp chat prompt from the 6.38 GB unsloth/gemma-3-270m-it-GGUF repository in 6.7 seconds with the image indexed versus 18.3 seconds for docker pull plus hf download, and read a single tensor from the 1.03 TB, 61-shard moonshotai/Kimi-K2-Instruct model in 3.4 seconds while moving 9.5 MB of the model. Range installs as a macOS or Linux release archive (x86-64 or arm64), requires Lima for its Linux VM on macOS, and needs root plus the nbd, erofs and overlay kernel modules on Linux.", "body_md": "README.md\n\nRange opens a shell in a container image, a Hugging Face repository, or an environment in S3 or on any HTTP server, without downloading it first. Only the bytes your program reads cross the network.\n\n``` bash\n$ range shell python:3.12\n$ range shell python:3.12 --mount hf://moonshotai/Kimi-K2-Instruct:/model\n$ range shell s3://<your-bucket>/dev.range\n```\n\nNo Docker, no daemon and no pull. Linux runs it natively. On macOS, Range runs Linux in a small VM that it manages itself.\n\nMeasured on EC2 in us‑east‑1, with the image indexed once. See bench.log.\n\ndemo.txt\n\n``` bash\n$ range run ghcr.io/ggml-org/llama.cpp:light-b11206 \\\n    --mount hf://unsloth/gemma-3-270m-it-GGUF:/model -- \\\n    llama-cli -m /model/gemma-3-270m-it-Q4_K_M.gguf -st \\\n    -p \"Why is the sky blue? Answer in one sentence.\"\n\nThe sky is blue because of a phenomenon called Rayleigh scattering,\nwhere blue light is scattered more than other colors.\n```\n\nRange opens the llama.cpp image from its registry and mounts the model repository at\n`/model`. The repository holds 6.38 GB in 24 files. Range reads one of them.\nWith the image indexed, the answer took 6.7 s from an empty cache.\n`docker pull` plus `hf download` took 18.3 s. The very first run,\nwhich also indexes the image, took 15.5 s.\n\n``` bash\n$ range run python:3.12 --mount hf://moonshotai/Kimi-K2-Instruct:/model -- \\\n    du -sh --apparent-size /model\n959G    /model\n```\n\nKimi K2 is 1.03 TB in 61 shards. A Python script inside read its config, the header of one shard and one tensor. That took 3.4 s with the image indexed, and moved 9.5 MB of the model. The other 60 shards never left Hugging Face.\n\nRange reads a file when a program opens it. A program that reads a whole model still downloads the whole model, once.\n\nbench.log\n\n| Empty cache to output | docker pull | Range, first run | Range, indexed | Range, again | \n|---|---|---|---|---|\n| Chat demo | 18.3 s 579 MB | 15.5 s 590 MB | 6.7 s 317 MB | 4.2 s 0 MB | \n| python:3.12 | 15.8 s 435 MB | 16.5 s 415 MB | 2.8 s 48 MB | 1.1 s 0 MB | \n| rust:1.82 | 19.6 s 569 MB | 22.4 s 546 MB | 8.0 s 125 MB | 1.6 s 0 MB | \n| eclipse-temurin:21 | 7.9 s 232 MB | 9.2 s 225 MB | 2.9 s 49 MB | 1.0 s 0 MB | \n| Kimi K2, 1 TB | not tried 1.03 TB | 17.6 s 433 MB | 3.4 s 56 MB | 1.2 s 0 MB | \n\nThe commands: import json and sqlite3, cargo --version, java -version, and a read of one Kimi K2 tensor. A first run reads each layer once to index it. The python:3.12 index is 4.1 MB. Medians of three, m6i.large, us‑east‑1, 28 September 2026. Every run starts empty, except \"again\". The bars replay at 3x speed.\n\nproblem.txt\n\nA machine that needs a large environment downloads all of it, every time, to use a small part.\n\nRange reads only the bytes each machine touches, and the next run fetches them before it asks.\n\ndesign.txt\n\n`ReadAt(offset, length) -> bytes`. Range turns a source into a disk, and turns\neach read of that disk into a ranged request to the source.\n\ninstall.txt\n\nA release archive for macOS or Linux, x86-64 or arm64. On macOS, Range also needs Lima for its Linux VM. Then open a shell in any image:\n\n``` bash\n$ curl -fsSL https://github.com/andreygrehov/range/releases/latest/download/range_$(uname -s)_$(uname -m).tar.gz | tar -xz\n$ brew install lima      # macOS only\n$ ./range shell python:3.12\n```\n\nFor your own environments, build once, and every first run reads lazily:\n\n``` bash\n$ range build --from-oci python:3.12 -o py.range\n$ range publish py.range s3://<your-bucket>/py.range\n$ range shell s3://<your-bucket>/py.range\n```\n\nA published artifact needs no indexing. The go1.23 demo artifact was ready in 0.41 s on its first run, and moved 6 MB of 1.03 GB.\n\nOr build from source, with Go 1.25 or newer:\n\n``` bash\n$ git clone https://github.com/andreygrehov/range && cd range && make install\n```\n\nLinux needs root and the nbd, erofs and overlay kernel modules.\n`range doctor` checks them. Windows works through WSL2, untested.\n\nabout_me.txt\n\nI am a software engineer at AWS. Range is my personal project.", "url": "https://wpnews.pro/news/range-open-a-1-tb-ai-model-in-3-seconds-without-downloading-it", "canonical_source": "https://getrange.sh/", "published_at": "2026-09-28 21:50:23+00:00", "updated_at": "2026-09-28 22:17:55.910529+00:00", "lang": "en", "topics": ["ai-infrastructure", "developer-tools", "ai-tools", "large-language-models"], "entities": ["Range", "Andrey Grehov", "AWS", "Hugging Face", "llama.cpp", "moonshotai/Kimi-K2-Instruct", "unsloth/gemma-3-270m-it-GGUF", "Lima"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/range-open-a-1-tb-ai-model-in-3-seconds-without-downloading-it", "markdown": "https://wpnews.pro/news/range-open-a-1-tb-ai-model-in-3-seconds-without-downloading-it.md", "text": "https://wpnews.pro/news/range-open-a-1-tb-ai-model-in-3-seconds-without-downloading-it.txt", "jsonld": "https://wpnews.pro/news/range-open-a-1-tb-ai-model-in-3-seconds-without-downloading-it.jsonld"}}