cd /news/large-language-models/inside-my-llama-cpp-setup-tuning-qwe… Β· home β€Ί topics β€Ί large-language-models β€Ί article
[ARTICLE Β· art-144232] src=dev.to β†— pub= topic=large-language-models verified=true sentiment=Β· neutral

Inside My llama.cpp Setup: Tuning Qwen 3.8 27B for 512K Context

A developer documented a llama.cpp configuration for running Unsloth's Qwen3.8-27B GGUF model at a 512K-token context window on an M5 MacBook Pro with 128 GB of unified memory, targeting parallel multi-agent coding workloads. The setup combines MTP speculative decoding with up to 8 draft tokens, full GPU layer offload (-ngl 99), YaRN RoPE scaling from a 262144-token base, f16 KV cache with offload, and 2 parallel slots. The writeup walks through each flag's purpose and the memory-versus-context trade-offs involved.

by read8 min views1 publishedOct 3, 2026

I've been tuning llama.cpp for local AI development, and the command line can quickly become a collection of cryptic flags.

Here's what my current configuration does, parameter by parameter.

I'm specifically focusing on maxing out the utilization of my system (which is MBP M5 with 128 GB Unified RAM), for multi-agent coding which requires parallel agents execution.

llama serve \
  -hf unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_XL \
  --spec-type draft-mtp \
  --spec-default \
  --spec-draft-n-max 8 \
  -ngl 99 \
  -c 524288 \
  --override-kv qwen2.context_length=int:524288 \
  --rope-scaling yarn \
  --yarn-orig-ctx 262144 \
  -b 16384 -ub 4096 \
  -t 16 \
  -tb 16 \
  -np 2 \
  -fa on \
  --cache-type-k f16 \
  --cache-type-v f16 \
  --kv-offload \
  --load-mode none \
  --host 127.0.0.1 \
  --port 8080

The easiest way to understand it is to divide the configuration into several areas:

-hf unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_XL

This tells llama.cpp to download and load the model from Hugging Face.

Breaking it down:

unsloth/ β€” Hugging Face repository ownerQwen3.8-27B β€” approximately 27 billion parametersGGUF β€” the model format used by llama.cpp UD-Q4_K_XL β€” the quantization Q4 means the model weights are approximately 4-bit quantized.

The trade-off is straightforward: lower precision produces a much smaller model and significantly reduces memory requirements, at some cost to numerical precision.

--spec-type draft-mtp

This enables speculative decoding using MTP (Multi-Token Prediction).

Instead of having the main model generate:

token β†’ token β†’ token β†’ token

the system uses a draft mechanism to propose multiple future tokens, which the main model then verifies.

Conceptually:

                 Draft model
                      β”‚
                      β–Ό
              token token token
                      β”‚
                      β–Ό
                Main model
                  verifies
                      β”‚
                      β–Ό
             accept several tokens

When several proposed tokens are accepted, generation can become substantially faster.

For this model, speculative decoding is one of the most important performance-related settings.

--spec-default

--spec-default

This enables the default speculative-decoding configuration associated with the selected speculation type.

In this case:

draft-mtp

It's generally not something I'd change unless I was experimenting with the underlying speculative-decoding implementation.

--spec-draft-n-max 8

This controls the maximum number of speculative tokens proposed ahead.

With:

8

the draft mechanism can attempt to predict up to eight tokens ahead.

Main model:
A

Draft:
A β†’ B β†’ C β†’ D β†’ E β†’ F β†’ G β†’ H

Main model verifies:
A B C D βœ“ βœ“ βœ“ βœ—

The more tokens you speculate, the greater the potential speedupβ€”but only if the draft predictions are good enough.

This is one of the parameters worth benchmarking:

2
4
8

Eight is an aggressive but reasonable value to test.

-ngl 99

This is short for:

--n-gpu-layers

It specifies how many model layers should be offloaded to the GPU.

99 effectively means:

Put as many layers as possible on the GPU.

It does not mean "use 99 GPU cores."

Think of it as:

CPU
 β”‚
 β”œβ”€β”€ some model layers
 β”‚
GPU
 β”‚
 └── most/all model layers

If the model fits comfortably on the GPU, -ngl 99 is generally what you want for performance.

-c 524288

This specifies the maximum context window.

The value is:

524,288 tokens = 512K tokens.

That is an enormous context window.

For comparison:

32K   = 32,768
128K  = 131,072
256K  = 262,144
512K  = 524,288

The important trade-off is that larger context requires more memory, particularly because of the KV cache.

For agentic coding workloads, however, having hundreds of thousands of tokens available can be extremely useful.

--override-kv qwen2.context_length=int:524288

This is different from -c.

You're overriding a value stored in the model's GGUF metadata:

qwen2.context_length

and setting it to:

524288

In other words, you're telling llama.cpp to treat the model as having a 512K context length.

This does not magically train the model for 512K context.

That's why the configuration also uses YaRN.

--rope-scaling yarn

This enables YaRN β€” Yet another RoPE extension.

RoPE stands for Rotary Position Embedding.

RoPE is part of how the transformer represents token positions:

token 1
token 2
token 3
...
token 262144

When extending the context beyond the model's original trained range, positional scaling is required.

YaRN provides a mechanism for extending that range.

In this configuration:

Original context:
262K

Target context:
524K

So the positional range is being extended by roughly 2Γ—.

--yarn-orig-ctx 262144

This tells YaRN:

The model's original context length is 262,144 tokens.

So the relevant configuration is:

Original:
262,144

Target:
524,288

Extension:
2Γ—

These two parameters work together:

--rope-scaling yarn
--yarn-orig-ctx 262144
-b 16384

This specifies the maximum number of tokens processed in a logical batch.

16,384 tokens

This primarily affects prompt processing / prefill.

For example, if you send a large prompt containing thousands of tokens, a larger batch can allow the GPU to process more tokens efficiently.

Larger batches can increase prompt-processing throughput, but they also consume more memory.

Importantly:

context = 524K
batch   = 16K

is perfectly valid.

The batch size does not limit the context window.

-ub 4096

This is the physical or micro-batch size.

It controls how many tokens are actually processed at one time.

The configuration therefore has:

Logical batch:
16,384

Physical batch:
4,096
16,384 tokens

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ 4,096 tokens     β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ 4,096 tokens     β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ 4,096 tokens     β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ 4,096 tokens     β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

This allows a large logical batch without requiring all 16K tokens to be processed simultaneously.

-ub is therefore particularly important for VRAM usage and prompt-processing performance.

-t 16

This specifies the number of CPU threads used for computation.

Here:

16 CPU threads

This does not mean 16 GPU cores.

How useful additional CPU threads are depends heavily on your CPU and on how much of the workload remains on the CPU.

-tb 16

This specifies the number of CPU threads used specifically for batch processing.

So the configuration is:

Normal computation: 16 threads
Batch computation:  16 threads

Whether 16 is optimal depends on your CPU.

If you're running on a high-core-count CPU, this is worth benchmarking.

More threads don't automatically mean higher performance.

-np 2

This enables two parallel sequences/requests.

                 Model
                   β”‚
          β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”
          β–Ό                 β–Ό
      Context #1         Context #2
      512K max           512K max

This is useful if you're running two concurrent requests or agents.

There is, however, a memory cost.

512K context
Γ—
2 parallel sequences

the potential KV-cache requirement becomes very large.

If you only ever run one request at a time, -np 1 may provide a better memory/performance balance.

-fa on

This enables Flash Attention.

Flash Attention is an optimized implementation of the attention mechanism designed to reduce memory traffic and improve performance.

It becomes particularly important at long context lengths.

For a 512K configuration, I'd keep:

-fa on
--cache-type-k f16

This specifies the datatype used for the Key portion of the KV cache.

You're using:

F16

or 16-bit floating point.

--cache-type-v f16

This specifies the datatype used for the Value portion of the KV cache.

So the current configuration is:

K = F16
V = F16

This provides high precision, but consumes considerably more memory than:

--cache-type-k q8_0
--cache-type-v q8_0

Given the combination of:

512K context
Γ—
2 parallel sequences
Γ—
F16 KV

this is one of the largest memory-consuming choices in the configuration.

--kv-offload

This tells llama.cpp to keep the KV cache on the GPU when possible.

That generally improves performance because it avoids repeatedly moving KV data between CPU and GPU.

The desired architecture for maximum performance is therefore approximately:

Model weights β†’ GPU
KV cache     β†’ GPU
Attention    β†’ GPU

assuming you have enough VRAM.

--load-mode none

This controls the model- mechanism.

none means that no special mode is being selected.

This isn't a setting I'd normally spend much time optimizing unless you're diagnosing model , memory mapping, or startup behavior.

--host 127.0.0.1

This makes the server listen only on the local machine.

So the server is accessible through:

127.0.0.1

but isn't directly exposed to other machines on the network.

This is a network/security setting, not an inference-performance setting.

--port 8080

The server listens on port:

8080

So your local API is effectively:

http://127.0.0.1:8080

This has essentially no impact on model performance.

Your command is essentially saying:

Run Qwen 3.8 27B using a Q4 quantization, put as much of the model as possible on the GPU, use MTP speculative decoding with up to eight speculative tokens, support a 512K context by extending the model's 256K positional range with YaRN, process prompts using 16K/4K batches, use 16 CPU threads, support two simultaneous sequences, use Flash Attention, keep the F16 KV cache on the GPU, and expose the model as a local HTTP server on port 8080.

The architecture looks roughly like this:

                    llama.cpp server
                           β”‚
               β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
               β”‚                       β”‚
          Request #1              Request #2
          512K max                512K max
               β”‚                       β”‚
               β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                           β”‚
                       KV Cache
                        F16/F16
                           β”‚
                    Flash Attention
                           β”‚
                  Qwen 3.8 27B Q4
                           β”‚
                     GPU (-ngl 99)
                           β”‚
                   MTP / speculative
                      decoding Γ—8
                           β”‚
                         Output

Not all parameters deserve equal attention.

--spec-draft-n-max 8
-c 524288
-np 2
--cache-type-k f16
--cache-type-v f16
-b 16384
-ub 4096
-fa on
-t 16
-tb 16
-ngl 99
--kv-offload
--override-kv
--rope-scaling
--yarn-orig-ctx
--load-mode
--host
--port

The three biggest trade-offs in this particular configuration are:

512K context  ↔  memory

F16 KV       ↔  memory / performance

16K / 4K batch ↔ VRAM / prompt throughput

And for generation speed, the most interesting parameter is probably:

MTP-8 ↔ speculative-token acceptance rate

The optimal configuration therefore isn't necessarily the one with the largest numbers. The goal is to find the point where GPU utilization, memory bandwidth, KV-cache size, batch size, and speculative-token acceptance work together rather than competing with each other.

── more in #large-language-models 4 stories Β· sorted by recency
── more on @llama.cpp 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/inside-my-llama-cpp-…] indexed:0 read:8min 2026-10-03 Β· β€”