14:37
2026-07-20
dev.to
large-language-models
Bonsai-27B on a Single 3090: What Works and What Doesn't
A developer tested the Bonsai-27B model on a single RTX 3090 using a Q4_K_M GGUF quant from prism-ml, achieving ~28 tok/s at 4K context and ~19 tok/s at 16K. The model excelled at structured extractioβ¦