cd/entity/LlamaBenchยท homeโ€บ entitiesโ€บ LlamaBench
grep -l @llamabench /news/*.json | wc -l โ†’ 1

LlamaBench

mentions 1 type Organization feed RSS

// recent coverage 1 mentions

10:03
2026-07-23
promptcube3.com
artificial-intelligence

DGX Spark Inference: 90 tok/s on Large Models

A new inference stack for DGX Spark clusters achieves 55-90 tok/s on large models without speculative decoding, according to internal tests by WoolyAI. The stack enables multi-model agentic workflows โ€ฆ

// co-occurs with top 5 entities
// topics top 4 topics