cd/entity/llama-cpp-pythonΒ· homeβ€Ί entitiesβ€Ί llama-cpp-python
grep -l @llama-cpp-python /news/*.json | wc -l β†’ 3

llama-cpp-python

mentions 3 type Organization feed RSS

// recent coverage 3 mentions

00:00
2026-08-20
nobodywho.ai
machine-learning

Use fewer threads for CPU inference

A benchmark of Gemma4-E4B on an AMD Ryzen 7040 CPU shows that using 8 threads instead of 16 improves prompt processing from 94.66 to 105.94 tokens per second and token generation from 11.64 to 16.15 t…

15:33
2026-06-16
dev.to
large-language-models

Can You Tell When an LLM API Swaps in a Cheaper Model?

A developer found that LLM API providers can swap in cheaper models undetected, and the intuitive method of flagging low-quality responses fails because cheaper models produce more predictable text wi…

// co-occurs with top 8 entities
// topics top 6 topics