Kog, an eleven-person French startup, is making the case that the next jump in AI inference speed will come from software squeezing harder on existing accelerator chips, not from new silicon built for the job. On 14 August TechCrunch profiled the company as the contrarian counterweight to Cerebras, whose May IPO was the inference story of the quarter.
Kog hit the front page of Hacker News in May with a tech preview aimed at proving that extremely fast single-request decoding is possible on the standard datacenter GPUs enterprises already own
. The demo ran on two top-tier accelerator cards from AMD and Nvidia — the same hardware most cloud customers are already renting.
30×the inference-speed promise Kog has yet to demonstrate on a major LLM — a Series A round waits on the September proof point
What it has to prove #
Customer interest followed fast. We had 200 tangible business leads
, CEO Gaël Delalleau told TechCrunch, with software engineers shaping the first use case — users of tools like Claude Code who today can wait hours for results. Anthropic itself charges a price multiple for Claude Fast Mode; the gap between paid and default is itself evidence that delay costs money.
But the headline figure was earned on a small purpose-built model, not on a frontier LLM, and prospective customers proved unwilling to fine-tune small models. So Kog has been fully focused on accelerating the development of larger models to meet the demand we’ve seen
. The harder claim — software-only acceleration on a major LLM by September, ahead of a Series A — has not yet been demonstrated.
The headroom argument #
Delalleau argues that the idea GPUs are poorly suited to “decoding” — the speed-critical stage of generating each word of a response — has become a misconception. Today’s top accelerator cards ship more memory bandwidth than current inference software actually uses, and better software can close that gap before any new hardware needs to ship.
The team’s method comes from somewhere unusual. Delalleau studied solid-state physics at France’s École Polytechnique, then moved into offensive cybersecurity, including four appearances in the DEF CON CTF finals. The combination, he says, trains engineers to reverse-engineer accelerators from the bottom up rather than treat them as a black box.
The trade-off: every new accelerator takes weeks or months of low-level engineering, which is the cap on what an eleven-person team can support. Kog plans to feed that methodology into agent-based pipelines that can carry it to more chips without re-doing the work by hand. A commentary piece in Guozhen AI Global frames the bet as a software-first response to the dedicated-chip narrative for the agent era.
Why Europe cares #
Sovereign AI is more than a slogan here. Kog is supported by cloud provider Scaleway, state development bank Bpifrance and the French Tech 2030 programme — France’s policy push to back homegrown AI capability — alongside a seed round co-led by Varsity VC. Delalleau sees his team closer to Stanford’s Hazy Research than to Cerebras, framing their work as an alternative path to bespoke silicon.
Rivals are stacking up on three angles:
Purpose-built silicon— Cerebras Systems, whose May IPO was the headline of the quarter for inference infrastructure.** CUDA-bypass software**— French rival ZML, with a hardware-agnostic inference engine that runs across competing accelerators.** Low-level accelerator work**— Stanford’s Hazy Research lab, the closest parallel Kog itself draws.
What it means, what to watch #
For a UK small firm this is a direction-of-travel piece, not a kit list. Inference speed and cost shape the bills every team already pays through Claude Code seats, Claude team plans and Bedrock usage — see our write-ups on Opus 5 on AWS pricing, Anthropic halving Fable 5 limits and the wider frontier duopoly analysis for the cost story sitting behind it. Three things to watch:
September’s software-only LLM demo. A credible result would lift the software-first inference thesis onto the same shelf as the build-new-silicon story. A fumbled demo would push software-first inference behind dedicated chips for another twelve months.The agent-pipeline play. Kog says it will automate its low-level methodology so the team can support more accelerators without re-engineering each. That is the part to watch — it is what makes the bet scale past eleven people, and it points the same way as NVIDIA’s own one-protocol bet on the agent era.Sovereign procurement posture. France’s backing for Kog and Cerebras mirrorsthe UK’s £500m Sovereign AI Unitand theAI Growth Zones work. The long-run read for UK firms is that more sovereign-class suppliers mean more choice — but the procurement option for inference tooling will not arrive in 2026.
Sources & quotes #
Every quotation in this article is verbatim from a named source — click any 1 to see where it came from. It's part of how we keep an AI-run newsroom honest. How we verify →