cd/entity/GGUF· home› entities› GGUF
grep -l @gguf /news/*.json | wc -l → 97

GGUF

mentions 97 type Organization page 2/5 feed RSS

// recent coverage 97 mentions

00:00
2026-09-16
nobodywho.ai
ai-infrastructure

NobodyWho vs Cactus: On-Device Inference Engine Comparison

NobodyWho Edge runs GGUF models through llama.cpp with no conversion step, while Cactus uses its own Cactus Quants (CQ) rotation-and-codebook quantization format from 4-bit down to 1-bit, according to…

16:00
2026-09-15
gladlabs.io
ai-infrastructure

Llama.cpp vs vLLM vs SGLang

Glad Labs decided not to switch its self-hosted inference stack from Ollama to vLLM after reviewing its own call logs, which showed only one to three concurrent calls at most against roughly 50 calls …

10:14
2026-09-14
dev.to
ai-tools

llama.cpp vs Ollama in 2026: Which Runtime Should You Run?

A technical comparison examines the tradeoffs between Ollama and llama.cpp for local LLM inference, framing the choice as one between a managed model service and a toolkit operated directly. The guide…

15:53
2026-09-11
blog.kilo.ai
large-language-models

How to Choose a Local LLM: Models, Hardware, and Quantization

Atomic Chat published a guest blog guide on selecting local large language models, recommending that users start with a GGUF Q4_K_M quantization if it fits their hardware. The guide provides memory-es…

09:09
2026-09-11
dev.to
large-language-models

KV Cache on 16 GB GPUs: Making Long Context Actually Fit

A developer published a practical guide to fitting long-context LLM inference on 16 GB GPUs by budgeting VRAM for the KV cache, which grows with every active token and sequence. The guide provides a K…

09:56
2026-09-08
anup.io
machine-learning

TIL: What’s actually different about GGUF and ONNX?

A technical comparison explains that GGUF stores model weights and metadata for llama.cpp, which implements the architecture itself, while ONNX stores a computation graph that any compatible runtime c…

00:00
2026-09-08
unite.ai
artificial-intelligence

How Long Before a Real Crackdown on AI Model Decensoring?

Open-source AI models are increasingly being stripped of their safety filters and redistributed at scale, echoing the warez scene of 1995–2010, according to an analysis on Unite.AI. The practice invol…

15:11
2026-09-07
gist.github.com
developer-tools

sentence piece model from gguf

A developer shared a Python script that extracts a SentencePiece tokenizer model from GGUF files, enabling the use of the tokenizer with the sentencepiece library. The approach reads token, score, and…

16:46
2026-09-01
promptcube3.com
artificial-intelligence

Stop chasing prompts and start building deterministic systems

AI engineering is shifting from prompt optimization to building deterministic systems with observability and local-first architectures, according to a technical article. The piece advocates for treati…

16:45
2026-08-29
promptcube3.com
large-language-models

Running massive LLMs on consumer hardware is a financial

Quantization and pruning techniques such as GPTQ, AWQ, GGUF, SparseGPT, and LoRA-based pruning enable running large language models on consumer hardware by reducing memory footprint, with 4-bit quanti…

11:26
2026-08-26
github.com
artificial-intelligence

Axera AX8850 LLM running ggufs

A custom llama.cpp backend (ggml-axcl) now runs Qwen3-0.6B directly from GGUF files on an Axera AX8850 NPU accelerator card (M5Stack LLM-8850) hosted on a Raspberry Pi 5, achieving 1.3-2.7 tokens per …

22:00
2026-08-23
fratepietro.com
artificial-intelligence

Ferrox on Metal: at parity with llama.cpp, and past it

Ferrox, a pure-Rust GGUF inference engine, now runs mixture-of-experts (MoE) prefill on Apple Metal 2.4x faster, reaching 1402 tok/s on OLMoE-1B-7B and closing the gap to llama.cpp from 2.62x behind t…

← prev page 2 / 5 next →
// co-occurs with top 8 entities
// topics top 6 topics