Bonsai-27B on a Single 3090: What Works and What Doesn't
A developer tested the Bonsai-27B model on a single RTX 3090 using a Q4_K_M GGUF quant from prism-ml, achieving ~28 tok/s at 4K context and ~19 tok/s at 16K. The model excelled at structured extractio…