Most explanations of llama.cpp start with the tool itself: download this, run that command, hereβs your chatbot. That skips the actual origin story, and the origin story is where the interesting engineering lives. So this piece starts one layer down, at the thing llama.cpp is built on: a small math library called GGML, and the file format that grew out of it. Only once that foundation is solid does it make sense to talk about llama.cpp itself, because by the end, youβll see llama.cpp for what it really is: a thin, purpose-built wrapper around GGML.
GGML stands for Georgi Gerganov Machine Learning, named after its creator. It began as, and still is, two things at once:
Weβll come back to the format later. First, the library.
Gerganov built the GGML framework because he wanted to run LLaMA on his Macβs CPU, with none of the GPU or heavy Python dependencies the official release assumed. GGML is what made that possible, and its design reflects that goal directly:
PyTorch tensors are built for training: dynamic shapes, autograd bookkeeping, dispatch across arbitrary backends, and eager execution by default (each op runs immediately, with allocation and deallocation cycles as tensors come and go β modern PyTorch softens this with caching allocators, but the underlying model is still fundamentally dynamic).
GGML tensors are built for inference on constrained hardware, and the struct itself shows it:
struct ggml_tensor { enum ggml_type type; // Data type (e.g., GGML_TYPE_F32, GGML_TYPE_Q4_0) struct ggml_backend_buffer * buffer; // Pointer to hardware buffer (CPU RAM or VRAM) int64_t ne[GGML_MAX_DIMS]; // Number of elements per dimension (shape) size_t nb[GGML_MAX_DIMS]; // Stride/bytes per dimension // Computation graph metadata enum ggml_op op; // The op that produced this tensor, if any int32_t flags; // GGML_TENSOR_FLAG_INPUT, _PARAM, etc. struct ggml_tensor * src[GGML_MAX_SRC]; // Input tensors to op // View optimization struct ggml_tensor * view_src; // Non-NULL if this is a slice/view of another tensor size_t view_offs; void * data; // Raw pointer to the weights char name[GGML_MAX_NAME]; // e.g. "blk.0.attn_v.weight"};
Two design choices stand out:
Because GGML builds the full computation graph before running anything, it knows in advance exactly how much memory every intermediate tensor needs. It reserves that memory once and reuses the same addresses across every token and every batch. It can even fuse some operations together for extra efficiency.
PyTorchβs eager mode, historically, worked the opposite way: allocate as you go, free when done, repeat for every operation in every forward pass. Repeated allocation and deallocation isnβt free β the allocator has to hunt for a suitably sized contiguous block each time. Modern PyTorch narrows this gap with caching allocators that hold onto memory between iterations, but GGMLβs graph-first approach sidesteps the problem entirely rather than mitigating it after the fact.
This isnβt just a theoretical claim β it shows up directly in a simple matrix multiplication benchmark. I ran the same 1024Γ1024 quantized weight matrix through both GGML (Q8_0 weights, quantized-on-the-fly activations) and PyTorch (quantize_dynamic, the same int8-weight-and-activation scheme), across two shapes: multiplying against a 1024Γ1024 batch, and multiplying against a single 1024Γ1 vector β the shape that actually occurs during one step of autoregressive token generation.
The two rows arenβt just different numbers, theyβre different bottlenecks entirely. At batch size 1024, every weight gets reused 1024 times before itβs evicted from cache, so the workload is compute-bound, and PyTorchβs mature, heavily-tuned kernels win. At batch size 1, each weight byte is read from memory and used exactly once β thereβs no reuse to amortize, so the workload becomes bound by memory bandwidth and by how much per-call overhead each framework carries. Thatβs where GGMLβs graph-once, allocate-once design pays off: thereβs no allocator to consult, no Python-to-C++ dispatch to cross, no tensor bookkeeping to redo, just the same pre-built graph running again with the same buffers.
That second row is the one that actually matters for local inference. Generating text one token at a time is a sequence of matrix-vector multiplies, not matrix-matrix ones, so the batch-size-1 result is the more representative predictor of how each framework performs during real chat-style generation, not the batch-size-1024 one.
(Numbers above are from a single machine, one CPU, one matrix size β the crossover point between βPyTorch winsβ and βGGML winsβ will shift with matrix size, thread count, and hardware. Worth keeping that caveat in mind before generalizing too far from one benchmark.)
This is the part thatβs easy to get wrong, so itβs worth being precise. βGGMLβ as a file format wasnβt one static thing, it was three formats, released in sequence, each patching a problem with the last:
In practice, if you go download an old model file today with βggmlβ in its filename, what youβre actually holding is almost always GGJT. Thatβs exactly what happened in the case study below.
When people say βGGML formatβ in the context of old llama.cpp models, they usually mean the old binary model-file format used by the GGML/llama.cpp ecosystem β not to be confused with the ggml_tensor C structure used by the GGML runtime itself. Those are two different things sharing one name.
The old format is essentially a binary serialization of a modelβs configuration, vocabulary, and tensors. Its structure is rigid: the reader has to already know what each field means and exactly where it appears in the file, since nothing in the file itself is self-describing.
At a high level, an old GGML/GGJT model file looks like:
GGML/GGJT fileββββ Headerβ βββ magicβ βββ versionββββ fixed hyperparametersββββ Vocabularyβ βββ token lengthβ βββ token bytesβ βββ token scoreββββ Tensors βββ tensor header βββ dimensions βββ tensor name βββ alignment padding βββ raw tensor data
A) Header
The header begins with a 4-byte magic number that identifies the file format, followed by a version field.
For a GGJT v3 file, the relevant layout is:
4 bytes β magic number4 bytes β version
B) Hyperparameter list
The seven hyperparameters are:
n_vocab β vocabulary sizen_embd β model embedding / hidden dimensionn_mult β architecture-specific feed-forward multiplen_head β number of attention headsn_layer β number of Transformer layersn_rot β RoPE dimensionftype β integer identifier for the model's tensor/quantization format
Since each hyperparameter occupies 4 bytes, the seven fields consume:
7 Γ 4 = 28 bytes
right after the initial 8-byte magic/version portion. The old format has a fixed binary header layout, and it reads fields strictly in that predetermined order.
C) Vocabulary
The vocabulary immediately follows the hyperparameter list. For each of the n_vocab entries, the reader consumes:
4 bytes β token lengthtoken length β token bytes4 bytes β tokenizer score (float32)
Conceptually, the layout in memory looks like:
[token length][token bytes][score][token length][token bytes][score][token length][token bytes][score]...
The token itself is stored as raw bytes rather than in a fixed-size field, so the parser reads the first 4 bytes to learn how many bytes the token occupies, then advances by exactly that amount.
The score belongs to the tokenizerβs vocabulary β it is not a neural-network weight, logit, or attention score.
D) Tensor data
After the vocabulary, tensors are stored one after another. Each tensor has a small binary descriptor followed by its actual data:
4 bytes β n_dims4 bytes β name_len4 bytes β dtype4 Γ n_dims β dimensionsname_len β tensor name0β31 bytes β alignment paddingn_bytes β raw tensor data
For example, a tensor might be:
name = "layers.0.attention.wq.weight"shape = (4096, 4096)dtype = Q4_0
The number of elements is:
n_elems = dimensions[0] Γ dimensions[1] Γ ...
For a quantized tensor, its storage size is determined by the quantization type:
n_bytes = n_elems Γ tysize / blksize
where:
blksize = number of elements represented by one quantization blocktysize = number of bytes occupied by that block
For Q4_0, each block represents a fixed number of weights and stores both the quantized values and the associated quantization scale. The tensor doesnβt separately parse those scales β it treats the whole block as raw tensor data. For standard Q4_0 quantization, blksize = 32 and tysize = 18 bytes.
Each quantization scheme fixes its own blksize/tysize. Here's how it's actually defined, straight from ggml-common.h:
typedef uint16_t ggml_half; #define QK4_0 32typedef struct { ggml_half d; // delta uint8_t qs[QK4_0 / 2]; // nibbles / quants} block_q4_0;static_assert(sizeof(block_q4_0) == sizeof(ggml_half) + QK4_0 / 2, "wrong q4_0 block size/padding");
used in ggml.c as:
static const struct ggml_type_traits type_traits[GGML_TYPE_COUNT] = { ... [GGML_TYPE_Q4_0] = { .type_name = "q4_0", .blck_size = QK4_0, .type_size = sizeof(block_q4_0), .is_quantized = true, .to_float = (ggml_to_float_t) dequantize_row_q4_0, .from_float_ref = (ggml_from_float_t) quantize_row_q4_0_ref, }, ...
In the Q4_0 scheme, the tensor data is organized as blksize = 32 and type_size = 16 + 2 bytes, meaning a single block holds 32 weights and occupies 18 bytes.
E) Alignment padding
Before the tensor data begins, the parser may skip some padding bytes so that the tensor data starts at an aligned address, typically a multiple of 32 bytes in the old format.
tensor metadata βpadding β32-byte-aligned tensor data
This alignment matters because it lets the tensor data be accessed efficiently, and itβs specifically what makes memory-mapped possible.
F) Important distinction: file format vs. GGML tensor types
The old file format tells the how to find and interpret tensors. The ggml_type enum tells it what kind of data each tensor actually contains:
GGML_TYPE_F32GGML_TYPE_F16GGML_TYPE_Q4_0GGML_TYPE_Q5_0GGML_TYPE_Q8_0GGML_TYPE_Q4_K...
For example:
dtype = 2 βGGML_TYPE_Q4_0
The number 2 here is a type identifier, not β2-bit quantization.β
I downloaded llama-2-7b.ggmlv3.q4_0.bin from a legacy TheBloke/Llama-2-7B-GGML repository. Its file size:
Size : 3,791,725,184 bytes (3.53 GiB)
The script used to decode the following info from the file is linked here.
The file begins as:
==========================================================================================HEADER (offset, raw bytes, decoded field)========================================================================================== [0x000000-0x000003] 74 6a 67 67 -> magic number : 'tjgg'[0x000004-0x000007] 03 00 00 00 -> version : 3
The first four bytes are the magic bytes 74 6a 67 67, which correspond to 'tjgg'. The bytes look reversed when read as ASCII because the integer magic value is stored in little-endian order β the important point is that these bytes identify the file as the GGJT variant.
The next four bytes, 03 00 00 00, decode as the integer version = 3.
The seven fixed hyperparameters follow:
[0x000008-0x00000b] 00 7d 00 00 β n_vocab : 32000 [0x00000c-0x00000f] 00 10 00 00 β n_embd : 4096 [0x000010-0x000013] 00 01 00 00 β n_mult : 256 [0x000014-0x000017] 20 00 00 00 β n_head : 32 [0x000018-0x00001b] 20 00 00 00 β n_layer : 32 [0x00001c-0x00001f] 80 00 00 00 β n_rot : 128 [0x000020-0x000023] 02 00 00 00 β ftype : 2
Which implies:
Vocabulary size = 32000Hidden dimension = 4096Attention heads = 32Transformer layers = 32RoPE dimension = 128Tensor format = MOSTLY_Q4_0
Thereβs also a useful consistency check:
n_embd / n_head = 4096 / 32 = 128
which matches n_rot = 128.
The important thing to notice is that these values are stored without descriptive key names anywhere in the file. The reader knows the first 4-byte integer means n_vocab, the next means n_embd, and so on purely because the format specification hard-codes that order.
The vocabulary begins immediately after the fixed hyperparameter list, at 0x000024. For each token, the parser performs:
read 4-byte token length βread that many token bytes to get the token value βread 4-byte float score βmove to the next token
For token 0:
[0x000024-0x000027] 05 00 00 00 β tok_len : 5 [0x000028-0x00002c] 20 e2 81 87 20 β token value : ' β ' [0x00002d-0x000030] 00 00 00 00 β score : 0.0
For token 3:
[0x000041-0x000044] 01 00 00 00 β tok_len : 1 [0x000045-0x000045] 00 β token : '\x00' [0x000046-0x000049] 00 00 00 00 β score : 0.0
The important observation is that the token length alone determines exactly how far the parser advances β thereβs no fixed-width field to fall back on.
A few more examples:
token [0] [0x000024-0x000027] 05 00 00 00 -> tok_len : 5 [0x000028-0x00002c] 20 e2 81 87 20 -> token : ' β ' [0x00002d-0x000030] 00 00 00 00 -> score : 0.0 token [1] [0x000031-0x000034] 00 00 00 00 -> tok_len : 0 [0x000035-0x000034] -> token : '' [0x000035-0x000038] 00 00 00 00 -> score : 0.0 token [2] [0x000039-0x00003c] 00 00 00 00 -> tok_len : 0 [0x00003d-0x00003c] -> token : '' [0x00003d-0x000040] 00 00 00 00 -> score : 0.0 token [3] [0x000041-0x000044] 01 00 00 00 -> tok_len : 1 [0x000045-0x000045] 00 -> token : '\x00' [0x000046-0x000049] 00 00 00 00 -> score : 0.0 token [4] [0x00004a-0x00004d] 01 00 00 00 -> tok_len : 1 [0x00004e-0x00004e] 01 -> token : '\x01' [0x00004f-0x000052] 00 00 00 00 -> score : 0.0 ... (31992 more tokens, not shown) ... token [31997] [0x0699c1-0x0699c4] 03 00 00 00 -> tok_len : 3 [0x0699c5-0x0699c7] e6 94 b6 -> token : 'ζΆ' [0x0699c8-0x0699cb] 00 f4 f7 c6 -> score : -31738.0 token [31998] [0x0699cc-0x0699cf] 03 00 00 00 -> tok_len : 3 [0x0699d0-0x0699d2] e5 bc 98 -> token : 'εΌ' [0x0699d3-0x0699d6] 00 f6 f7 c6 -> score : -31739.0 token [31999] [0x0699d7-0x0699da] 03 00 00 00 -> tok_len : 3 [0x0699db-0x0699dd] e7 bb 99 -> token : 'η»' [0x0699de-0x0699e1] 00 f8 f7 c6 -> score : -31740.0
After processing all 32,000 entries, the parser reaches the first tensor.
Example tensor entries from the file:
==========================================================================================TENSORS (offset, raw bytes, decoded field)========================================================================================== tensor [0] [0x0699e2-0x0699e5] 02 00 00 00 -> n_dims : 2 [0x0699e6-0x0699e9] 15 00 00 00 -> name_len: 21 [0x0699ea-0x0699ed] 02 00 00 00 -> dtype : 2 [0x0699ee-0x0699f5] 00 10 00 00 00 7d 00 00 -> dims : (4096, 32000) [0x0699f6-0x069a0a] 74 6f 6b 5f 65 6d 62 65 64 64 69 6e 67 73 2e 77 ... -> name : 'tok_embeddings.weight' [0x069a0b-0x069a1f] 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 ... -> padding : 21 bytes (aligns next field to 32-byte offset) [0x069a20-0x069a31] 1b 00 89 77 95 9d a9 e5 aa 69 04 67 b5 75 68 63 89 9e -> data : first 18 of 73,728,000 raw bytes (Q4_0, 4,096,000 block(s) of 32 elements) ... (291 more tensors, not shown) ...
The first tensor begins at 0x0699e2. Its descriptor reads: n_dims = 2, name_len = 21 bytes, dtype = 2 (Q4_0).
The two dimensions are (4096, 32000), giving:
n_elems = 4096 Γ 32000 = 131,072,000
The parser then encounters 21 bytes of padding, moving the tensor data to the next aligned address, 0x069a20. The actual tensor data begins there.
Verifying the tensor size against the quantization scheme (Q4_0):
Total elements = 4096 Γ 32000 = 131,072,000 individual weightsTotal blocks = 131,072,000 / 32 = 4,096,000Total size = 4,096,000 Γ 18 = 73,728,000 bytes
So every tensor in Q4_0 follows the same pattern: 2 bytes of scale, followed by 16 bytes of quantized weights, repeated per block. For this first tensor, block 0 breaks down as:
(Iβm deliberately not decoding these specific bytes into actual float values here β that depends on endianness and exactly how the scale and 4-bit values combine, which is really a Part 3 topic once quantization itself is properly explained.)
The key structural point:
metadata βpadding βraw Q4_0 blocks
The Q4_0 scales live inside those raw blocks β the GGML parser never creates a separate βscale section.β Every other tensor in the file is decoded the same way.
The full structure of this file, condensed:
GGJT v3 Llama-2-7Bββββ Headerβ βββ magicβ βββ versionβ βββ n_vocab = 32000β βββ n_embd = 4096β βββ n_mult = 256β βββ n_head = 32β βββ n_layer = 32β βββ n_rot = 128β βββ ftype = 2 β Q4_0ββββ Vocabularyβ βββ token 0β βββ token 1β βββ ...β βββ token 31999ββββ Tensors βββ tok_embeddings.weight β βββ descriptor β βββ padding β βββ Q4_0 data β βββ norm.weight β βββ descriptor β βββ padding β βββ F32 data β βββ output.weight β βββ descriptor β βββ padding β βββ Q4_0 data β βββ layers.0.attention.wq.weight β βββ descriptor β βββ padding β βββ Q4_0 data β βββ layers.0.attention.wk.weight β βββ descriptor β βββ padding β βββ Q4_0 data β βββ ... 286 more tensors
The important idea here is that the file is, fundamentally, just a sequential binary stream. The parser maintains a running offset and repeats:
read field βadvance offset βread next field βadvance offset β...
GGJTβs header is a fixed list of untyped hyperparameters: seven fields, in a fixed order, each an opaque 4-byte integer. Thatβs fine as long as every model is a LLaMA, but by mid-2023 llama.cpp needed to support Mistral, Falcon, and a growing list of architectures, each with different config needs: rotary embedding bases, grouped-query attention parameters, sliding window sizes, and more.
None of that fits into βseven fixed integers.β Every time the project needed to store one more piece of information, the choices were: break the format, or bolt on a hack. Multiplied across dozens of contributors and architectures, that became unsustainable. GGJT had no concept of a typed, extensible metadata dictionary β everything had to be known and hardcoded by the in advance.
Thatβs the exact problem GGUF was designed to solve.
GGUF keeps GGJTβs core idea β a single, self-contained, mmap-friendly file β but replaces the fixed hyperparameter list with a typed key-value metadata store. Any number of arbitrary, named, typed fields can be added without breaking older readers.
Version history (each version is a small, deliberate change, not a rewrite):
struct gguf_file_t { gguf_header_t header; gguf_tensor_info_t tensor_infos[header.tensor_count]; uint8_t _padding[]; // pad to nearest multiple of ALIGNMENT uint8_t tensor_data[]; // raw weights};
Header β 24 bytes, fixed size:
struct gguf_header_t { uint32_t magic; // "GGUF" -> 0x47 0x47 0x55 0x46 uint32_t version; uint64_t tensor_count; uint64_t metadata_kv_count; gguf_metadata_kv_t metadata_kv[metadata_kv_count];};
Each metadata entry β a typed key/value pair:
struct gguf_metadata_kv_t { gguf_string_t key; // hierarchical, e.g. "general.architecture" gguf_metadata_value_type value_type; gguf_metadata_value_t value;};
Each tensorβs descriptor β separate from its data:
struct gguf_tensor_info_t { gguf_string_t name; // e.g. "blk.0.attn_q.weight", max 64 bytes uint32_t n_dimensions; // currently at most 4 uint64_t dimensions[n_dimensions]; ggml_type type; uint64_t offset; // byte offset into tensor_data, must be ALIGNMENT-aligned};
The alignment itself is just another metadata key (general.alignment, default 32) β even that low-level detail is extensible rather than hardcoded.
Unlike the older GGML/GGJT model-file formats, GGUF was designed to be extensible and self-describing. Instead of storing a fixed list of hyperparameters in a predetermined order, GGUF stores model metadata as typed key-value (KV) pairs.
At a high level, a GGUF file looks like:
GGUF fileββββ Headerβ βββ magicβ βββ versionβ βββ tensor countβ βββ metadata KV countββββ Metadataβ βββ key-value pairsββββ Tensor Informationβ βββ descriptor for every tensorββββ Alignment paddingββββ Tensor Data βββ actual F32/F16/quantized weight bytes
The GGUF header is 24 bytes:
4 bytes β magic4 bytes β version8 bytes β tensor count8 bytes β metadata KV count
The four fields tell the how to interpret the rest of the file:
Unlike the old GGML format, the header itself doesnβt contain a fixed list like n_vocab, n_embd, n_layer. Those values now live in the metadata section as named KV pairs.
The main improvement in GGUF is its typed metadata system. Each metadata entry conceptually contains:
ββββββββββββββββββββββββββββββββ Key length ββββββββββββββββββββββββββββββββ€β Key ββββββββββββββββββββββββββββββββ€β Value type ββββββββββββββββββββββββββββββββ€β Value ββββββββββββββββββββββββββββββββ
The important difference from the old GGML format is that metadata is now named and typed. Instead of the reader assuming:
first integer β n_vocabsecond integer β n_embdthird integer β n_mult...
GGUF explicitly tells the reader:
"llama.embedding_length" β UINT32 β 4096 KEY VALUE TYPE VALUE
This makes the format much easier to extend. New architecture-specific parameters can be added as new KV entries without redesigning the header at all. Unlike the fixed-size header, the metadata sectionβs size is dynamic β it depends entirely on how many KV pairs there are and how large their keys and values are.
This section doesnβt contain the actual model weights β it contains a descriptor for every tensor, telling the how to find and interpret that tensorβs data.
Conceptually, each tensor descriptor contains: name, n_dimensions, dimensions, type, offset.
The important field here is offset. It tells the where the tensor's data begins, relative to the start of the tensor-data blob β not relative to the start of the file. This separation matters: the tensor-information section works like a directory or index, while the tensor-data section holds the actual weights.
The final major section holds the actual numerical tensor data. For a quantized tensor such as Q4_0, the raw bytes contain the complete quantized representation, including whatever the quantization format needs (like scales).
The tensor descriptor tells the runtime:
What is this data? βname + shape + type + offset βWhere is it? βoffset into tensor-data blob βRead the corresponding raw bytes
Before the tensor-data blob begins, GGUF can insert alignment padding so the data region starts aligned, typically to a 32-byte boundary.
The GGUF counterpart of the same model, llama-2-7b.Q4_0.gguf (downloaded from β TheBloke/Llama-2-7B-GGUF):
Size : 3,825,807,040 bytes (3.56 GiB)
The script used to decode the following info from the file is linked here..
The first 24 bytes are:
==========================================================================================HEADER (offset, raw bytes, decoded field)========================================================================================== [0x000000-0x000003] 47 47 55 46 β magic : GGUF [0x000004-0x000007] 02 00 00 00 β version : 2 [0x000008-0x00000f] 23 01 00 00 00 00 00 00 β tensor_count : 291 [0x000010-0x000017] 13 00 00 00 00 00 00 00 β metadata_kv_count : 19
So immediately, before reading anything else, the already knows:
This is a GGUF file βGGUF version = 2 β291 tensors are described later β19 metadata KV pairs must be read
The metadata starts immediately after the 24-byte header:
==========================================================================================METADATA (19 key-value pairs) (offset, raw bytes, decoded field)========================================================================================== kv [0] [0x000018-0x00001f] 14 00 00 00 00 00 00 00 -> key len : 20 [0x000020-0x000033] 67 65 6e 65 72 61 6c 2e 61 72 63 68 69 74 65 63 ... -> key : 'general.architecture' [0x000034-0x000037] 08 00 00 00 -> value_type : 8 (value_type 8 = STRING) [0x000038-0x00003f] 05 00 00 00 00 00 00 00 -> value len : 5 [0x000040-0x000044] 6c 6c 61 6d 61 -> value : 'llama' kv [1] [0x000045-0x00004c] 0c 00 00 00 00 00 00 00 -> key len : 12 [0x00004d-0x000058] 67 65 6e 65 72 61 6c 2e 6e 61 6d 65 -> key : 'general.name' [0x000059-0x00005c] 08 00 00 00 -> value_type : 8 (value_type 8 = STRING) [0x00005d-0x000064] 08 00 00 00 00 00 00 00 -> value len : 8 [0x000065-0x00006c] 4c 4c 61 4d 41 20 76 32 -> value : 'LLaMA v2' kv [2] [0x00006d-0x000074] 14 00 00 00 00 00 00 00 -> key len : 20 [0x000075-0x000088] 6c 6c 61 6d 61 2e 63 6f 6e 74 65 78 74 5f 6c 65 ... -> key : 'llama.context_length' [0x000089-0x00008c] 04 00 00 00 -> value_type : 4 (value_type 4 = UINT32) [0x00008d-0x000090] 00 10 00 00 -> value : 4096 kv [3] [0x000091-0x000098] 16 00 00 00 00 00 00 00 -> key len : 22 [0x000099-0x0000ae] 6c 6c 61 6d 61 2e 65 6d 62 65 64 64 69 6e 67 5f ... -> key : 'llama.embedding_length' [0x0000af-0x0000b2] 04 00 00 00 -> value_type : 4 (value_type 4 = UINT32) [0x0000b3-0x0000b6] 00 10 00 00 -> value : 4096 kv [4] [0x0000b7-0x0000be] 11 00 00 00 00 00 00 00 -> key len : 17 [0x0000bf-0x0000cf] 6c 6c 61 6d 61 2e 62 6c 6f 63 6b 5f 63 6f 75 6e ... -> key : 'llama.block_count' [0x0000d0-0x0000d3] 04 00 00 00 -> value_type : 4 (value_type 4 = UINT32) [0x0000d4-0x0000d7] 20 00 00 00 -> value : 32 kv [5] [0x0000d8-0x0000df] 19 00 00 00 00 00 00 00 -> key len : 25 [0x0000e0-0x0000f8] 6c 6c 61 6d 61 2e 66 65 65 64 5f 66 6f 72 77 61 ... -> key : 'llama.feed_forward_length' [0x0000f9-0x0000fc] 04 00 00 00 -> value_type : 4 (value_type 4 = UINT32) [0x0000fd-0x000100] 00 2b 00 00 -> value : 11008 kv [6] [0x000101-0x000108] 1a 00 00 00 00 00 00 00 -> key len : 26 [0x000109-0x000122] 6c 6c 61 6d 61 2e 72 6f 70 65 2e 64 69 6d 65 6e ... -> key : 'llama.rope.dimension_count' [0x000123-0x000126] 04 00 00 00 -> value_type : 4 (value_type 4 = UINT32) [0x000127-0x00012a] 80 00 00 00 -> value : 128 kv [7] [0x00012b-0x000132] 1a 00 00 00 00 00 00 00 -> key len : 26 [0x000133-0x00014c] 6c 6c 61 6d 61 2e 61 74 74 65 6e 74 69 6f 6e 2e ... -> key : 'llama.attention.head_count' [0x00014d-0x000150] 04 00 00 00 -> value_type : 4 (value_type 4 = UINT32) [0x000151-0x000154] 20 00 00 00 -> value : 32 ... (11 more KV pairs, not shown) ...
There are 19 KV pairs of metadata total. Taking the first one apart field by field:
[0x000018-0x00001f] β key length : 20[0x000020-0x000033] β key : 'general.architecture'[0x000034-0x000037] β value_type : 8 (STRING)[0x000038-0x00003f] β value length: 5[0x000040-0x000044] β value : 'llama'
So the raw bytes decode to: general.architecture β STRING β "llama". The next KV pair is general.name β STRING β "LLaMA v2", then llama.context_length β UINT32 β 4096, and so on through the remaining pairs.
==========================================================================================TENSOR INFO (291 tensors) (offset, raw bytes, decoded field)========================================================================================== Note: this is just the DESCRIPTOR list -- name, shape, type, and offset into the tensor_data blob. The raw weights themselves live later in the file, all together (see next section). tensor [0] [0x0b0b29-0x0b0b30] 11 00 00 00 00 00 00 00 -> name len : 17 [0x0b0b31-0x0b0b41] 74 6f 6b 65 6e 5f 65 6d 62 64 2e 77 65 69 67 68 ... -> name : 'token_embd.weight' [0x0b0b42-0x0b0b45] 02 00 00 00 -> n_dims : 2 [0x0b0b46-0x0b0b55] 00 10 00 00 00 00 00 00 00 7d 00 00 00 00 00 00 -> dims : (4096, 32000) [0x0b0b56-0x0b0b59] 02 00 00 00 -> type : 2 [0x0b0b5a-0x0b0b61] 00 00 00 00 00 00 00 00 -> offset : 0 (type 2 = Q4_0, 131,072,000 elements, ~73,728,000 bytes of tensor data at data-blob offset 0) tensor [1] [0x0b0b62-0x0b0b69] 16 00 00 00 00 00 00 00 -> name len : 22 [0x0b0b6a-0x0b0b7f] 62 6c 6b 2e 30 2e 61 74 74 6e 5f 6e 6f 72 6d 2e ... -> name : 'blk.0.attn_norm.weight' [0x0b0b80-0x0b0b83] 01 00 00 00 -> n_dims : 1 [0x0b0b84-0x0b0b8b] 00 10 00 00 00 00 00 00 -> dims : (4096,) [0x0b0b8c-0x0b0b8f] 00 00 00 00 -> type : 0 [0x0b0b90-0x0b0b97] 00 00 65 04 00 00 00 00 -> offset : 73728000 (type 0 = F32, 4,096 elements, ~16,384 bytes of tensor data at data-blob offset 73,728,000) tensor [2] [0x0b0b98-0x0b0b9f] 15 00 00 00 00 00 00 00 -> name len : 21 [0x0b0ba0-0x0b0bb4] 62 6c 6b 2e 30 2e 66 66 6e 5f 64 6f 77 6e 2e 77 ... -> name : 'blk.0.ffn_down.weight' [0x0b0bb5-0x0b0bb8] 02 00 00 00 -> n_dims : 2 [0x0b0bb9-0x0b0bc8] 00 2b 00 00 00 00 00 00 00 10 00 00 00 00 00 00 -> dims : (11008, 4096) [0x0b0bc9-0x0b0bcc] 02 00 00 00 -> type : 2 [0x0b0bcd-0x0b0bd4] 00 40 65 04 00 00 00 00 -> offset : 73744384 (type 2 = Q4_0, 45,088,768 elements, ~25,362,432 bytes of tensor data at data-blob offset 73,744,384) tensor [3] [0x0b0bd5-0x0b0bdc] 15 00 00 00 00 00 00 00 -> name len : 21 [0x0b0bdd-0x0b0bf1] 62 6c 6b 2e 30 2e 66 66 6e 5f 67 61 74 65 2e 77 ... -> name : 'blk.0.ffn_gate.weight' [0x0b0bf2-0x0b0bf5] 02 00 00 00 -> n_dims : 2 [0x0b0bf6-0x0b0c05] 00 10 00 00 00 00 00 00 00 2b 00 00 00 00 00 00 -> dims : (4096, 11008) [0x0b0c06-0x0b0c09] 02 00 00 00 -> type : 2 [0x0b0c0a-0x0b0c11] 00 40 e8 05 00 00 00 00 -> offset : 99106816 (type 2 = Q4_0, 45,088,768 elements, ~25,362,432 bytes of tensor data at data-blob offset 99,106,816) tensor [4] [0x0b0c12-0x0b0c19] 13 00 00 00 00 00 00 00 -> name len : 19 [0x0b0c1a-0x0b0c2c] 62 6c 6b 2e 30 2e 66 66 6e 5f 75 70 2e 77 65 69 ... -> name : 'blk.0.ffn_up.weight' [0x0b0c2d-0x0b0c30] 02 00 00 00 -> n_dims : 2 [0x0b0c31-0x0b0c40] 00 10 00 00 00 00 00 00 00 2b 00 00 00 00 00 00 -> dims : (4096, 11008) [0x0b0c41-0x0b0c44] 02 00 00 00 -> type : 2 [0x0b0c45-0x0b0c4c] 00 40 6b 07 00 00 00 00 -> offset : 124469248 (type 2 = Q4_0, 45,088,768 elements, ~25,362,432 bytes of tensor data at data-blob offset 124,469,248) tensor [5] [0x0b0c4d-0x0b0c54] 15 00 00 00 00 00 00 00 -> name len : 21 [0x0b0c55-0x0b0c69] 62 6c 6b 2e 30 2e 66 66 6e 5f 6e 6f 72 6d 2e 77 ... -> name : 'blk.0.ffn_norm.weight' [0x0b0c6a-0x0b0c6d] 01 00 00 00 -> n_dims : 1 [0x0b0c6e-0x0b0c75] 00 10 00 00 00 00 00 00 -> dims : (4096,) [0x0b0c76-0x0b0c79] 00 00 00 00 -> type : 0 [0x0b0c7a-0x0b0c81] 00 40 ee 08 00 00 00 00 -> offset : 149831680 (type 0 = F32, 4,096 elements, ~16,384 bytes of tensor data at data-blob offset 149,831,680) tensor [6] [0x0b0c82-0x0b0c89] 13 00 00 00 00 00 00 00 -> name len : 19 [0x0b0c8a-0x0b0c9c] 62 6c 6b 2e 30 2e 61 74 74 6e 5f 6b 2e 77 65 69 ... -> name : 'blk.0.attn_k.weight' [0x0b0c9d-0x0b0ca0] 02 00 00 00 -> n_dims : 2 [0x0b0ca1-0x0b0cb0] 00 10 00 00 00 00 00 00 00 10 00 00 00 00 00 00 -> dims : (4096, 4096) [0x0b0cb1-0x0b0cb4] 02 00 00 00 -> type : 2 [0x0b0cb5-0x0b0cbc] 00 80 ee 08 00 00 00 00 -> offset : 149848064 (type 2 = Q4_0, 16,777,216 elements, ~9,437,184 bytes of tensor data at data-blob offset 149,848,064) tensor [7] [0x0b0cbd-0x0b0cc4] 18 00 00 00 00 00 00 00 -> name len : 24 [0x0b0cc5-0x0b0cdc] 62 6c 6b 2e 30 2e 61 74 74 6e 5f 6f 75 74 70 75 ... -> name : 'blk.0.attn_output.weight' [0x0b0cdd-0x0b0ce0] 02 00 00 00 -> n_dims : 2 [0x0b0ce1-0x0b0cf0] 00 10 00 00 00 00 00 00 00 10 00 00 00 00 00 00 -> dims : (4096, 4096) [0x0b0cf1-0x0b0cf4] 02 00 00 00 -> type : 2 [0x0b0cf5-0x0b0cfc] 00 80 7e 09 00 00 00 00 -> offset : 159285248 (type 2 = Q4_0, 16,777,216 elements, ~9,437,184 bytes of tensor data at data-blob offset 159,285,248) ... (283 more tensor descriptors, not shown) ...
The tensor-information section describes all 291 tensors. Taking the first one apart:
[0x0b0b29-0x0b0b30] β name length : 17[0x0b0b31-0x0b0b41] β name : 'token_embd.weight'[0x0b0b42-0x0b0b45] β n_dims : 2[0x0b0b46-0x0b0b55] β dimensions : (4096, 32000)[0x0b0b56-0x0b0b59] β type : 2 = Q4_0[0x0b0b5a-0x0b0b61] β offset : 0
This tensor holds 4096 Γ 32000 = 131,072,000 elements, and since it's Q4_0, its data occupies roughly 73,728,000 bytes.
The next descriptor, tensor [1], reads:
name : blk.0.attn_norm.weightn_dims : 1dimensions : (4096,)type : 0 = F32offset : 73,728,000
This tells the runtime that blk.0.attn_norm.weight's data begins 73,728,000 bytes into the tensor-data blob, and since it's F32, it needs 4096 Γ 4 = 16,384 bytes.
So the tensor-information section is, in effect, a directory: it tells the runtime what every tensor is and exactly where to find its data.
==========================================================================================TENSOR DATA (offset, raw bytes, decoded field)========================================================================================== [0x0b4eaf-0x0b4ebf] 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 ... -> padding : 17 bytes (aligns tensor data blob to 32-byte offset) [0x0b4ec0-0x0b4ecf] 1b 00 89 77 95 9d a9 e5 aa 69 04 67 b5 75 68 63 -> token_embd.weight: first 16 of 73,728,000 bytes (Q4_0) [0x4704ec0-0x4704ecf] 00 00 f3 3c 00 00 5f 3c 00 00 01 3b 00 00 47 3c -> blk.0.attn_norm.weight: first 16 of 16,384 bytes (F32) [0x4708ec0-0x4708ecf] d0 1d 58 65 d9 d6 6c cb b8 4a 6e b9 85 79 96 07 -> blk.0.ffn_down.weight: first 16 of 25,362,432 bytes (Q4_0)
After all 291 tensor descriptors, the file reaches the actual tensor-data region. Thereβs first some alignment padding at [0x0b4eaf-0x0b4ebf] (17 bytes), and the tensor-data blob itself starts at absolute file offset 0x0b4ec0.
This distinction matters: the offsets stored in the tensor descriptors are relative to the start of this blob, not absolute file offsets.
For token_embd.weight (descriptor offset = 0), the absolute file location is:
tensor_data_start + tensor_offset = 0x0b4ec0 + 0 = 0x0b4ec0
Its first bytes, 1b 00 89 77 95 9d a9 e5 aa 69 04 67 b5 75 68 63, are the beginning of the raw Q4_0 representation of token_embd.weight.
For blk.0.attn_norm.weight, the descriptor says offset = 73,728,000, so its data begins at tensor_data_start + 73,728,000, and the first bytes there β 00 00 f3 3c 00 00 5f 3c 00 00 01 3b 00 00 47 3c β are ordinary 32-bit floats, since this tensor is F32.
For blk.0.ffn_down.weight, offset = 73,744,384, so its data begins at tensor_data_start + 73,744,384, and the bytes there are Q4_0 data, laid out exactly as described earlier: 2 bytes of scale followed by the packed quantized weights.
GGUF Llama-2 fileββββ HEADERβ ββ βββ magic = "GGUF"β βββ version = 2β βββ tensor_count = 291β βββ metadata count = 19ββββ METADATAβ ββ βββ general.architecture β "llama"β βββ general.name β "LLaMA v2"β βββ llama.context_length β 4096β βββ llama.embedding_length β 4096β βββ llama.block_count β 32β βββ llama.feed_forward_length β 11008β βββ llama.rope.dimension_count β 128β βββ llama.attention.head_count β 32β βββ ... 11 more KV pairsββββ TENSOR INFORMATIONβ ββ βββ token_embd.weightβ β shape = (4096, 32000)β β type = Q4_0β β offset = 0β ββ βββ blk.0.attn_norm.weightβ β shape = (4096,)β β type = F32β β offset = 73,728,000β ββ βββ blk.0.ffn_down.weightβ β shape = (11008, 4096)β β type = Q4_0β β offset = 73,744,384β ββ βββ blk.0.ffn_gate.weightβ βββ blk.0.ffn_up.weightβ βββ blk.0.ffn_norm.weightβ βββ blk.0.attn_k.weightβ βββ blk.0.attn_output.weightβ βββ ... 283 more tensor descriptorsββββ ALIGNMENT PADDINGβ βββ 17 bytesββββ TENSOR DATA β βββ token_embd.weight β Q4_0 raw bytes β βββ blk.0.attn_norm.weight β F32 raw bytes β βββ blk.0.ffn_down.weight β Q4_0 raw bytes β βββ blk.0.ffn_gate.weight β Q4_0 raw bytes β βββ ... remaining tensor data
Same model, same q4_0 quantization, yet the GGUF file is about 34 MB bigger (3,825,807,040 vs. 3,791,725,184 bytes). The quantized tensor data itself is unchanged; what grew is the metadata. GGJT's header held seven raw integers and a bare list of (token, score) pairs. GGUF's key-value store carries the same information plus a lot more that GGJT had no room for: explicit architecture identifiers, human-readable names and licensing fields, full tokenizer metadata (merges, token types, special-token IDs), and per-model fields specific to the architecture. That richness is precisely the extensibility GGJT couldn't offer, and it costs a few extra megabytes to store.
Everything discussed above, GGML the tensor library, and GGML/GGMF/GGJT/GGUF the file formats, is the engine and the fuel tank. None of it has any idea what a βchat templateβ is, what temperature or top-p sampling mean, or how to turn a stream of token IDs into readable text. Thatβs not GGMLβs job.
llama.cpp is the layer that turns raw tensor math into an actual LLM you can talk to. Itβs the orchestration code sitting on top of GGML: it knows the shape of a transformer forward pass, handles tokenization, applies prompt templates, manages the KV cache and context window, and runs the sampling logic (temperature, top-p, repetition penalties) that turns raw logits into the next token. None of that is GGMLβs concern, GGML just multiplies matrices fast and hands back numbers.
Put plainly:
This is also why GGUF isnβt a βllama.cpp format,β itβs the GGML projectβs format. Thatβs why other GGML-based tools, like whisper.cpp, can read GGUF files too: the format belongs to the engine, not to any one application built on top of it.
Thatβs the full arc: GGML the library, GGML the format (and its GGMF/GGJT detour), how GGUF fixed what GGJT couldnβt extend, and where llama.cpp actually sits relative to all of it, a wrapper, not the engine. Part 2 goes one level deeper into the ecosystem around the engine: what .safetensors is (the format Hugging Face exposes most LLMs in today) and how llama.cpp converts that, along with LoRA adapters, into .gguf.
Dissecting llama.cpp, Part 1: From GGML to GGUF, and Why llama.cpp Is Just the Wrapper was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.