cd/entity/GGUF· home› entities› GGUF
grep -l @gguf /news/*.json | wc -l → 97

GGUF

mentions 97 type Organization page 1/5 feed RSS

// recent coverage 97 mentions

08:59
2026-10-01
dev.to
large-language-models

When Chat Templates Go Wrong

A developer analysis explains that local language models which load cleanly but over-generate, ignore system prompts, or drop earlier instructions are usually suffering from a mismatched or misrendere…

16:38
2026-09-30
dev.to
large-language-models

Testing the 9x Smaller Local LLM Claim on a GPU-less VPS

A developer tested Prism's ternary-quantized Bonsai 2 27B model, a 27-billion-parameter LLM compressed into a single GGUF file under 6 GB, on a GPU-less Hetzner VPS using Prism's llama.cpp fork. The C…

19:09
2026-09-28
returneditor.ai
ai-tools

Why we removed Ollama from Return (and what replaced it)

Return, a document analysis tool for lawyers, replaced Ollama with a bundled llama-server from llama.cpp as a sidecar process after Ollama 0.12 introduced cloud models in September 2025 that proxy req…

15:04
2026-09-28
gist.github.com
large-language-models

Qwen 3.8 Flash Next config LLama.cpp

A developer benchmarked the Qwen3.8-Flash-Next GGUF quant (AD-4.27bpw Q4_K_M, 94.5 GB across 33 shards) on a single RTX 5060 Ti 16GB with 64GB DDR4, publishing llama.cpp server configurations that sus…

22:00
2026-09-27
fratepietro.com
artificial-intelligence

Frink Inference: the Rust alternative to llama.cpp

Frink, a pure-Rust inference engine for GGUF models created by developer Antonello F, now runs 100 architectures on its generic path with evidence, each verified against llama.cpp's own logits via lib…

07:18
2026-09-23
dev.to
large-language-models

How to Run Local LLMs on Apple Silicon

A developer guide details how to run local large language models on Apple Silicon, noting that as of Ollama 0.19 (March 2026), Ollama replaced its Metal-backed llama.cpp inference path with Apple's ML…

07:16
2026-09-23
dev.to
large-language-models

LLM Quantization Explained for Mac Users

A developer published an explainer on LLM quantization for Mac users, detailing how weight precision reduction works and why GGUF filenames like Q4_K_M, Q5_K_M, and Q4_0 are not interchangeable at the…

00:00
2026-09-23
mindstudio.ai
artificial-intelligence

Bonsai 2 27B: A 27B Model That Runs in Under 6GB on a Laptop

Prism ML released Bonsai 2 27B, a ternary-quantized version of Qwen3.8-27B that stores weights as -1, 0, or +1 and fits in roughly 6 to 8.6GB while retaining 98.2% of the original FP16 model's benchma…

03:06
2026-09-22
gist.github.com
large-language-models

Qwen 3.8 27B on RTX 5090 at 90-120tps

A developer published a llama-server configuration that runs a Qwen 3.8 27B NVFP4 model with MTP speculative decoding on an RTX 5090, reporting throughput of 90-120 tokens per second. The setup uses a…

00:00
2026-09-22
huggingface.co
large-language-models

Transformers now runs llama.cpp quants

Hugging Face added support for running GGUF models in its transformers library, letting users load llama.cpp-format checkpoints via `from_pretrained` and generate locally, with initial focus on Apple …

00:00
2026-09-18
mindstudio.ai
large-language-models

Run Bonsai 2 27B Locally on Apple Silicon With MLX

Bonsai 2 27B, a 27-billion-parameter ternary language model built on the Qwen3.8-27B hybrid-attention architecture, ships at 8.60GB total through the MLX 2-bit pack — a 7.67GB ternary language model p…

22:00
2026-09-17
fratepietro.com
large-language-models

A 27B model in 5.5 GB: Bonsai ternary on Ferrox, locally

Ferrox v0.23.0, released today, adds support for PrismML's Ternary-Bonsai-2-27B, a 27-billion-parameter model quantized to 1.75 bits per weight that occupies 5.5 GB on disk and runs on a 16 GB laptop,…

21:28
2026-09-17
discuss.huggingface.co
ai-tools

Ollama pull Error invalid model name

Ollama's model-name parser rejects remote Hugging Face references whose model component exceeds 80 characters, according to a technical analysis of the parser in Ollama's types/model/name.go file. A C…

20:03
2026-09-16
dev.to
large-language-models

A 4 GB Laptop GPU Beats a 12-Core CPU by 4.3x on Gemma 4

A developer benchmarked llama.cpp serving Google's Gemma 4 E2B Q4_0 GGUF on a single laptop, comparing CPU-only inference against the machine's 4 GB GTX 1650 Ti Max-Q, with the two arms differing only…

page 1 / 5 next →
// co-occurs with top 8 entities
// topics top 6 topics