cd/entity/llama-cpp-python· home› entities› llama-cpp-python
grep -l @llama-cpp-python /news/*.json | wc -l → 5

llama-cpp-python

mentions 5 type Organization feed RSS

// recent coverage 5 mentions

08:32
2026-09-23
til.simonwillison.net
large-language-models

Using Llama-cpp-Python grammars to generate JSON

Llama.cpp added grammar-constrained output generation on August 17, letting developers restrict a large language model's next-token selection so responses exactly match a specified grammar, and the ll…

00:00
2026-09-22
nobodywho.ai
large-language-models

Jev in 25 lines of Python

NobodyWho published a parody blog post on September 22, 2026 showing a Jev-style decision model implemented in 25 lines of Python using the Qwen3-0.6B-GGUF model via llama-cpp-python, which classifies…

00:00
2026-08-20
nobodywho.ai
machine-learning

Use fewer threads for CPU inference

A benchmark of Gemma4-E4B on an AMD Ryzen 7040 CPU shows that using 8 threads instead of 16 improves prompt processing from 94.66 to 105.94 tokens per second and token generation from 11.64 to 16.15 t…

15:33
2026-06-16
dev.to
large-language-models

Can You Tell When an LLM API Swaps in a Cheaper Model?

A developer found that LLM API providers can swap in cheaper models undetected, and the intuitive method of flagging low-quality responses fails because cheaper models produce more predictable text wi…

// co-occurs with top 8 entities
// topics top 6 topics