22:00
2026-08-23
fratepietro.com
artificial-intelligence
Ferrox on Metal: at parity with llama.cpp, and past it
Ferrox, a pure-Rust GGUF inference engine, now runs mixture-of-experts (MoE) prefill on Apple Metal 2.4x faster, reaching 1402 tok/s on OLMoE-1B-7B and closing the gap to llama.cpp from 2.62x behind t…