# Benchmarking Gemma 4 E2B on CPU with llama.cpp: A Practical Local AI Experiment

> Source: <https://dev.to/hardyweb/benchmarking-gemma-4-e2b-on-cpu-with-llamacpp-a-practical-local-ai-experiment-56nb>
> Published: 2026-10-09 00:02:22+00:00

I want to run a local language model for everyday office work without depending entirely on a cloud-based AI service.

My intended use cases include:

The challenge is that not every computer has the same CPU. Some machines have older processors with only a few cores, while others have newer CPUs with more cores and threads.

Instead of assuming that one configuration works everywhere, I want to establish a baseline and use the same benchmarking method across different machines.

This article records my initial experiments with **Gemma 4 E2B using `llama-server` from llama.cpp**.

The initial experiment was performed on a resource-constrained machine.

| Component | Specification | 
|---|---|
| CPU | AMD Athlon 3000G | 
| CPU architecture | 2 physical cores, 4 logical threads | 
| Graphics | Radeon Vega Graphics | 
| Runtime environment | WSL | 
| WSL memory limit | 5 GB | 
| Inference backend | llama.cpp `llama-server` | 
| Inference mode | CPU-only | 
| Parallel slots | 1 | 

The main model file was:

```
gemma-4-E2B-it-qat-UD-Q4_K_XL.gguf
```

The model is an instruction-tuned Gemma 4 E2B GGUF using the QAT Q4_K_XL quantization variant.

The objective is not to establish a universal performance figure for this model. It is to understand how configuration choices affect inference on this particular machine.

*Note: The exact llama.cpp version/build identifier was not recorded in this initial benchmark. Future tests should record it because performance and available options can change between builds.*

I started with a conservative configuration:

```
llama-server \
  -m ./gemma-4-E2B-it-qat-UD-Q4_K_XL.gguf \
  -t 2 \
  -tb 2 \
  -c 2048 \
  -b 64 \
  -ub 16 \
  --host 127.0.0.1 \
  --port 5001
```

The initial reported average generation speed was approximately **4.6 tokens per second**.

I then tested different thread counts and batch sizes. The following table records the results observed during the experiments.

| Test | Main configuration change | Reported speed | 
|---|---|---|
| A | `-t 2 -tb 2 -c 2048 -b 64 -ub 16` | 4.6 t/s | 
| B | `-t 3 -tb 4 -c 2048 -b 64 -ub 16` | 5.9 t/s | 
| B, repeat | Same thread configuration | 6.1 t/s | 
| C | `-b 32 -ub 8` , with`-t 3 -tb 4` | 6.0 t/s | 
| D | `-t 4 -tb 4` , with the original batch sizes | 5.8 t/s | 

These results suggested that using three generation threads and four batch-processing threads was worth investigating further on this CPU.

Increasing the thread count did not automatically improve performance. The four-thread test was slower than the three-thread result in these particular trials.

However, these were exploratory measurements rather than a controlled benchmark. They should not be interpreted as proof that three threads will always be optimal.

Next, I added several parameters:

```
-np 1
-ngl 0
--temp 0.7
--top-p 0.8
--top-k 20
--metrics
--jinja
--cache-type-k q8_0
--cache-type-v q8_0
--reasoning-format none
```

The reported speed was approximately **6.4 tokens per second**.

Because several options were added together, this result does not tell us which individual parameter, if any, improved generation speed.

I then tested Flash Attention explicitly:

```
-fa on
```

The reported speed for that trial was approximately **6.2 tokens per second**.

This trial did not demonstrate an improvement, so I left Flash Attention out of the preferred configuration for now. A more controlled test would be needed to establish whether it helps on another CPU or build.

`-np 1`: Configures one parallel sequence slot for this single-user experiment.`-ngl 0`: Keeps model inference on the CPU rather than offloading model layers to a GPU.`--cache-type-k q8_0` and `--cache-type-v q8_0`: Select quantized data types for the key and value caches.`--metrics`: Enables the server's metrics endpoint.`--jinja`: Enables the relevant Jinja chat-template handling.`--temp`, `--top-p`, and `--top-k`: Control sampling behaviour, not a guaranteed performance improvement.`--reasoning-format none`: Configures reasoning-format handling for the server.
The available options and their precise behaviour should be checked against the [official llama.cpp server documentation](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md).

The strongest reported result from the exploratory trials was approximately **6.6 tokens per second**.

The configuration was:

```
llama-server \
  -m ./gemma-4-E2B-it-qat-UD-Q4_K_XL.gguf \
  -t 3 \
  -tb 4 \
  -c 2048 \
  -b 128 \
  -ub 32 \
  -np 1 \
  -ngl 0 \
  --temp 0.7 \
  --top-p 0.8 \
  --top-k 20 \
  --metrics \
  --jinja \
  --cache-type-k q8_0 \
  --cache-type-v q8_0 \
  --reasoning-format none \
  --host 127.0.0.1 \
  --port 5001
```

This is my **provisional baseline**, not a claim that the configuration is universally optimal.

The larger batch settings were present in the best reported trial. More repeatable testing is required to establish whether they improve performance consistently, particularly because batch settings can affect prompt processing differently from token generation.

The `--no-mmap` option has not been benchmarked in this experiment and should not be considered part of the tested configuration.

A later trial used a Malay prompt requesting a formal memorandum about periodic maintenance for internally developed systems using Laravel and Debian Linux.

The generated memorandum covered:

The reported statistics were:

| Metric | Result | 
|---|---|
| Prompt tokens | 126 | 
| Generated tokens | 391 | 
| Reported average generation speed | 6.0 t/s | 
| Approximate generation time from token count and speed | 65 seconds | 

The time is an estimate calculated from the reported token count and speed. It is not a direct end-to-end timing measurement.

The memorandum had a recognisable formal structure and its recommendations were generally relevant to the requested topic.

However, it also generated a specific date, 26 May 2024, even though the prompt did not provide that date. This is an important reminder that local language models can introduce unsupported details.

For actual office use, the application should provide authoritative information such as dates, recipient details, reference numbers, and departmental names. Generated letters and technical recommendations should still be reviewed before they are issued or acted upon.

The experiment therefore evaluates two separate dimensions:

A higher tokens-per-second figure does not necessarily mean a better office assistant.

The measurements above were collected during interactive experiments. The prompts and output lengths were not identical in every trial, and some trials changed several parameters at once.

Consequently, the results are useful for narrowing down configurations, but not for drawing definitive conclusions about individual options.

A more reliable benchmark should:

The llama.cpp server supports a Prometheus-compatible metrics endpoint when `--metrics` is enabled. Its reported metrics can help distinguish prompt-processing throughput from generation throughput. See the [server metrics documentation](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md).

The long-term objective is to repeat this experiment across computers with different CPU generations, core counts, and memory limits.

For each machine, I will record:

| Field | What to record | 
|---|---|
| Machine ID | A simple label, such as `PC-01` | 
| CPU | Exact processor model | 
| CPU topology | Physical cores and logical threads | 
| Memory | Installed RAM and the actual limit available to the inference environment | 
| Operating environment | Linux, WSL, or another supported environment | 
| llama.cpp build | Version, build details, and relevant compilation options | 
| Model | Exact model filename and quantization | 
| Context size | Value of `-c` | 
| Threads | Values of `-t` and`-tb` | 
| Batch settings | Values of `-b` and`-ub` | 
| Cache settings | K and V cache types | 
| Prompt and output | Prompt tokens and generated tokens | 
| Performance | Generation tokens/s and prompt tokens/s, when available | 
| Resource use | Peak RAM, CPU utilisation, swapping, and temperature where available | 
| Quality | Accuracy, language quality, formatting, and unsupported claims | 

Start with the provisional baseline and change one variable at a time.

**Stage 1 — Thread configuration**

Compare suitable values for `-t` and `-tb` based on the CPU's topology. For a CPU with two physical cores and four logical threads, for example, test a small range of thread settings instead of assuming that all logical threads will always help.

**Stage 2 — Context size**

Compare `-c 1024` with `-c 2048` while keeping other parameters constant. A smaller context may reduce memory requirements, but it also limits how much conversation or source material can fit into the context.

The `-c 1024` configuration is a proposed future test, not a completed measurement in this experiment.

**Stage 3 — Batch sizes**

Test `-b` and `-ub` independently where practical. Record prompt-processing throughput as well as generation speed, because the effect may differ between these workloads.

**Stage 4 — KV cache**

Only after establishing a stable baseline, consider comparing `q8_0` with other supported cache types, such as `q4_0`. Measure memory use, speed, stability, and output quality. Do not assume that a smaller cache will necessarily make inference faster.

**Stage 5 — Practical workload**

Run the same set of tasks on each machine:

This will help determine which configurations are suitable for actual use rather than merely optimising a single benchmark prompt.

The initial experiment suggests several useful lessons:

This experiment establishes a starting point for running Gemma 4 E2B with llama.cpp on modest CPU hardware.

The current provisional configuration uses three generation threads, four batch threads, a context size of 2048, and batch settings of 128 and 32. It achieved a reported best speed of approximately 6.6 tokens per second in the exploratory trials.

The next goal is not simply to make one computer faster. It is to build a repeatable method for finding practical configurations across several computers, including older CPUs and newer machines with more cores.

The long-term measure of success is a useful local assistant for everyday office work: responsive enough for conversation, capable of producing good first drafts, and reliable enough to support documentation and basic data analysis with appropriate human review.

**Benchmark status:** Initial exploratory phase completed. Further parameter testing is paused until a new machine or a controlled test session is selected.
