# A weekend with TensorFold on a MacBook: the engine mattered, the quant did not

> Source: <https://dev.to/defilan/a-weekend-with-tensorfold-on-a-macbook-the-engine-mattered-the-quant-did-not-394p>
> Published: 2026-09-28 17:38:49+00:00

A 27B dense model on my MacBook decodes at about 26 tokens a second. That is fine for chat and painful for an agent that writes a few thousand tokens per turn across a hundred turns. So when a new engine showed up claiming several times that on the same laptop, with output that is exactly the same as plain decoding, I dropped what I had planned for the weekend. I wanted to know two things. Was it real, and could LLMKube serve it to the rest of my cluster like any other model? By Sunday the answer to both was yes, and I had a list of what broke on the way there. This is that list, in order.

Almost none of the fast parts are mine. [TensorFold](https://github.com/ashhart/TensorFold) is Ash's work (MIT), and it had shipped five tagged releases in the day before I started. The DFlash2 draft model is z-lab's. The MLX 4-bit weights I tested came from Vontra. What I added is the part LLMKube exists for: getting a Mac-only engine behind a Kubernetes Service, so a Mac is one more inference node in a cluster that also has NVIDIA and AMD boxes in it.

Before trusting any claim I run it myself on my own hardware, with the author's own benchmark. The README had M5 Max numbers for Nemotron 3.5 Lightning and for Qwen3.8-27B with a DFlash2 drafter. The box is a MacBook Pro with an M5 Max and 128 GB of unified memory.

Nemotron matched the README closely, and one cell came in higher. Qwen3.8-27B is the one I care about, because it is a dense model the size people actually run for coding. On a Kubernetes YAML prompt it decoded at 220 tokens a second, where mlx_lm manages about 31.

| Qwen3.8-27B, greedy, 256 tokens | code | K8s YAML | prose | 
|---|---|---|---|
| llama-server (Q4_K_M), no speculation | 26.4 | 25.4 | 26.0 | 
| llama-server, MTP draft, n-max 3 | 50.8 | 56.0 | 39.9 | 
| mlx_lm.server | 31.4 | 31.5 | 31.2 | 
| **TensorFold + DFlash2** | **139** | **220** | **85** | 

Speed means nothing if the text changes, so I checked the exactness claim too. Two models, sampled and greedy, four kinds of prompt: in all 16 cases the drafted output matched the undrafted output byte for byte. One honest caveat from later in the week. My "four requests in flight still match" result ran on a version that queues concurrent requests, so it proves queueing is clean, not that batched verification is. Ash has since shipped real multi-stream support, and rerunning that is next on my list.

Notice what the drafts do per workload. Kubernetes YAML and code are repetitive and predictable, so drafts land most of the time. Prose is not, and the gain drops to about 2.7x. Agentic coding traffic is almost all code, YAML and tool calls. That is close to the best case for this engine.

My first theory was that DFlash2 was doing the work, and any engine with a good drafter would get there. llama.cpp recently added a DFlash draft mode, and z-lab publishes a GGUF of the same drafter, so I could test that directly: same target weights, same drafter, llama-server instead of TensorFold.

| Same DFlash2 drafter | code | K8s YAML | prose | 
|---|---|---|---|
| llama-server, DFlash2, n-max 3 | 48.1 | 52.4 | 34.4 | 
| llama-server, DFlash2, n-max 7 (model card) | 46.8 | 58.7 | 24.6 | 
| TensorFold, DFlash2 | 139 | 220 | 85 | 

The theory was wrong. With the same drafter, llama.cpp lands at MTP speed on code and YAML, and at the recommended draft length it is slower than no speculation at all on prose. TensorFold gets about 2.5 to 3.7x more out of the identical draft model. As far as I can tell from reading its code, the difference is in verification. It checks a whole tree of candidate tokens in one pass, and its batched verify costs close to a single-row step. That is what I want to understand properly, and I will need a profiler to do it.

Metal cannot run inside a container, so LLMKube reaches a Mac through a small agent. The agent runs natively under launchd, starts engines on the Mac, and registers each one in the cluster as a Service with an EndpointSlice pointing at the Mac. Pods talk to a normal ClusterIP and never know the model is on a laptop. To teach the agent a new engine, I first had to restart it with some test flags.

The restart killed the coding model I run on this Mac every day. Homebrew had upgraded llama.cpp to 0.5.0 the day before, and 0.5.0 removed the `--mlock` flag. The agent always passes it. The old llama-server process had been started before the upgrade and kept running, so nothing looked wrong until something restarted it. Any reboot would have done the same.

It got worse from there. llama-server exited within a second on the bad flag, but the agent did not notice the process was gone. It kept polling a health endpoint for the full two-minute timeout, then retried, and the agent handles one event at a time. Every other model on that Mac sat in a queue behind a process that was already dead. The only log line was "timeout waiting for health check".

Two fixes came out of that night. The agent now asks llama-server which flag it supports and spells mlock either way. It also watches its child process, so a startup failure returns in about 150 ms with the exit status and the last lines of the engine's own log. A failure no longer takes two minutes and blocks everyone else.

Next I tried to pick an engine per model. The InferenceService has a `runtime` field for exactly that, and the agent reads it first. So I asked for `runtime: mlx-server` on a Mac model, and the API server rejected it as an unsupported value.

The history was worse than the error. The CRD's runtime enum only listed the in-cluster engines, and it defaulted the field to `llamacpp`. In June the agent was changed to prefer the per-model field over its own flag. Because admission filled in that default on every object, the flag never won again. Every Mac engine except llama.cpp had been unreachable from a normal InferenceService for three months. I checked all 25 services on my cluster, and every one had `runtime: llamacpp` stored.

The fix adds the Mac engines to the enum and drops the default. It also makes the operator refuse a Mac-only engine on a model that is not on a Mac, instead of quietly building a llama.cpp Deployment with the wrong label. While testing it I found two more gaps. An endpoint for a model the agent could not start kept saying `ready: true`, so Service traffic went to a closed port. And a model with an absolute local path as its source got treated as a download URL. Both are fixed too.

With runtime selection working again, TensorFold became one more executor in the agent. LLMKube does not install or pin TensorFold. You install a pinned release yourself, and the agent launches it. The whole thing is a spec change:

```
apiVersion: inference.llmkube.dev/v1alpha1
kind: Model
metadata:
  name: qwen38-27b-mlx
spec:
  source: lmstudio-community/Qwen3.8-27B-MLX-4bit
  format: mlx
  hardware:
    accelerator: metal
---
apiVersion: inference.llmkube.dev/v1alpha1
kind: InferenceService
metadata:
  name: qwen38-27b-tensorfold
spec:
  modelRef: qwen38-27b-mlx
  runtime: tensorfold
  replicas: 1
  contextSize: 65536
```

The service was Ready nine seconds after I applied it. I then measured from a pod inside the cluster, through the Service, which is how every real client will reach it. Kubernetes YAML ran at 227.5 tokens a second, code at 144.6 and prose at 89.9. Going direct to the Mac gave 231, 147 and 92, so the Kubernetes path costs about 2 percent. The output was byte-identical there too.

There was one more roadblock before the agents could use it, and it is a good one to know about. My first agent runs all died with `dial tcp ... i/o timeout`. The agent had registered the Mac under its Tailscale address. One of my nodes could reach that address, but the node where the coding agents run could not. My smoke test had passed only because its pod happened to land on the node that could. The agent's `--host-ip` flag fixed it with the Mac's LAN address.

The other half of the weekend was supposed to produce something only my lab could make: a quant calibrated on my own agent traffic. I scrubbed about 1.5 million tokens of real coding-agent sessions on LLMKube into a calibration set. Then I quantized Qwen3.8-27B from the BF16 weights with my set and with the generic set most community quants use.

My set won, barely, on agent text and lost on prose. At Q5_K_S even the agent edge was gone. The imatrix also saturated at around 65 thousand tokens, so a bigger corpus bought nothing. Custom calibration text does not pay for this model, and the generic set is fine.

Then I spent a day chasing a ghost. Even an 8-bit quant disagreed badly with BF16 on about 1 percent of agent tokens, which made no sense. It was not special-token handling, and it was not a Metal kernel, because the CPU backend was worse. MLX on the original weights showed the same thing, and so did Hugging Face transformers on a DGX Spark, running the reference recurrence in CUDA. So it is the model itself. On column-aligned listings like `ls -la`, Qwen3.8-27B becomes confidently wrong: it copies earlier file names into the wrong column. BF16 was lost on that text, so any quant looked bad against it. Once I stripped listings from the eval and switched to robust statistics, the numbers made sense again.

The next result looked like a win. Keeping the attention tensors at 8 bits cut divergence by about a quarter for 1.8 GB more. Then I ran the control I should have run first: spend the same bytes uniformly. Plain Q5_K_S, at a similar size, beat every targeted recipe, including bigger ones, by about 2x. The targeted "win" was just more bits.

| Qwen3.8-27B GGUF | size | agent text, top token matches BF16 | prose, top token matches BF16 | 
|---|---|---|---|
| Q4_K_M | 16.8 GB | 90.7% | 94.9% | 
| Attention + DeltaNet at 6-bit | 18.4 GB | 92.5% | 95.9% | 
| Q5_K_S | 19.0 GB | 94.1% | 97.1% | 
| Attention + DeltaNet at 8-bit | 20.1 GB | 92.8% | 96.0% | 

The MLX formats told a similar story. At the same size, MLX 4-bit trails GGUF Q4_K_M by a clear margin, and it takes about 5 bits in MLX to match it. The format TensorFold runs, 4-bit affine with group 64, agreed with BF16 on 86 percent of agent tokens against Q4_K_M's 91. On paper, the fast engine was running the least faithful weights in the test.

Token-level divergence is a proxy. What I care about is whether the model still writes working code. So I ran every problem in HumanEval+ and MBPP+, 542 of them, through Q4_K_M, Q5_K_S and TensorFold's 4-bit weights. The model-written code ran in a locked-down pod on the cluster, not on my laptop.

| EvalPlus, 542 problems | Q4_K_M | Q5_K_S | TensorFold 4-bit | 
|---|---|---|---|
| pass@1, plus tests | 81.2% | 81.2% | 82.1% | 
| generation wall-clock | 68 min | 94 min | 10.6 min | 

No detectable difference. A paired significance test on the problems that flipped gave p values between 0.47 and 1.0. The 3 to 4x gaps in divergence did not show up as failed problems. TensorFold did the whole set 6.4x faster.

EvalPlus problems are short and single-turn, though, and my real workload is an agent working an issue for an hour. So I gave three real LLMKube issues to the coding agents that already run on the cluster, once per contender. That made nine runs: the same agent and harness each time, with only the model endpoint changed. The agents are LLMKube's most demanding tenant, which makes them a good test of the platform too. I scored every branch with the same gate: the package tests, the full test suite, lint on two platforms, and a bite check. The bite check puts back the original code and requires the new tests to fail.

All nine passed, and all nine had tests that bite. TensorFold finished its three in 60 minutes of wall-clock. The two llama.cpp quants took about 185 minutes each. They shared the GPU for part of the run, which inflates their time, so read that as roughly 3x, not a precise figure.

Then I reviewed all nine by hand, because a green gate is not a review. Two of the nine had real defects the gate could not see. One draft hardcoded a test port, so its tests failed whenever something else on the machine was using it. Another sized a relative file path against the agent's working directory. Each quant ended up writing the version I picked for one issue, and those three are open as pull requests [#1928](https://github.com/defilantech/LLMKube/pull/1928), [#1929](https://github.com/defilantech/LLMKube/pull/1929) and [#1930](https://github.com/defilantech/LLMKube/pull/1930). Three issues is a small sample, and I would not claim the quants are equal on harder work. What I can say is that on real issues in this repo, the choice of quant did not change the outcome, and the choice of engine changed the wait by about 3x.

`uv tool install` line and a version check.`--tensorfold-bin` and set `runtime: tensorfold` on the InferenceService.`config/samples/inferenceservice_qwen38_27b_tensorfold.yaml` is the one above.`contextSize`.`--host-ip` to an address every node can reach,`--mlock`, which 0.5.0 removed.
The TensorFold runtime and every agent fix are on LLMKube main now and ship in the next release. Until then, build the agent from main.

TensorFold is pinned at v0.3.4.1 (`bb4b4a3`) with MLX 0.31.2. llama.cpp is the Homebrew 0.5.0 build `7fe450e19`, used for serving, conversion and every GGUF quant. Qwen3.8-27B is Apache-2.0 and TensorFold is MIT. The pull requests, in the order they landed:

If you run TensorFold on a Mac and your numbers disagree with mine, I would like to know. The [Discord](https://discord.gg/Ktz85RFHDv) is open.

*Originally posted on [llmkube.com](https://llmkube.com/blog/tensorfold-m5-max-engine-not-quant). LLMKube is Apache 2.0.*
