09:00
2026-08-05
fratepietro.com
artificial-intelligence
Building a Rust Inference Engine That Matches Llama.cpp
Developer Antonello F. released Ferrox, a pure-Rust inference engine that runs open LLMs locally on CPU, Apple Metal, or CUDA, achieving performance parity with llama.cpp on an Apple M2 Pro: 26.9 tok/โฆ