14:50
2026-08-14
techcrunch.com
artificial-intelligence
Kog is going deeper to squeeze more inference out of GPUs
French startup Kog claims its software can deliver 30x faster LLM inference on standard datacenter GPUs, demoing 3,000 tokens per second on a 2-billion-parameter model using AMD MI300X and NVIDIA H200…