cd/entity/llama 3.3B 70BΒ· homeβ€Ί entitiesβ€Ί llama 3.3B 70B
grep -l @llama 3.3b 70b /news/*.json | wc -l β†’ 1

llama 3.3B 70B

mentions 1 type Person feed RSS

// recent coverage 1 mentions

16:09
2026-08-27
forum.level1techs.com
large-language-models

So, what local inference models are we using?

A user reports running Gemma4:26b on a GPU at 2000+ tokens/s pre-fill and 80+ tokens/s evaluation, and Laguna S 2.1 118B on a CPU server at 13-14 tokens/s, expressing frustration at the lack of models…

// co-occurs with top 4 entities
// topics top 3 topics