Filling in the blanks: Inference using Phobos Developer Joa released Phobos v0.1.0, an LLM inference stack that compiles every kernel at runtime from the Phobos language, generating 328 tokens per second versus llama.cpp's 260 on a 0.8B model and 55 versus 31 on a 35B mixture-of-experts model on an RTX 2080 SUPER. Phobos lost to PrismML's newest llama.cpp fork on a ternary 27B model, decoding 45 versus 49 tokens per second, though it processed prompts 1.27 times faster, and retained three quarters of its decode speed when the desktop reclaimed 2 GiB mid-session against llama.cpp's 30%. Prompt processing on small models still trails llama.cpp at 60 to 80% of its speed. TLDR: Phobos, the tiny kernel language from my last post