cd/entity/FerroxΒ· homeβ€Ί entitiesβ€Ί Ferrox
grep -l @ferrox /news/*.json | wc -l β†’ 5

Ferrox

mentions 5 type Organization feed RSS

// recent coverage 5 mentions

07:52
2026-09-19
medium.com
large-language-models

Bonsai: A 27B reasoning model on a 16 GB M2 Mac, with Ferrox

A developer has released Bonsai, a 27-billion-parameter reasoning model designed to run on a 16 GB M2 Mac, alongside a companion tool called Ferrox. The project targets local inference of a large reas…

22:00
2026-09-17
fratepietro.com
large-language-models

A 27B model in 5.5 GB: Bonsai ternary on Ferrox, locally

Ferrox v0.23.0, released today, adds support for PrismML's Ternary-Bonsai-2-27B, a 27-billion-parameter model quantized to 1.75 bits per weight that occupies 5.5 GB on disk and runs on a 16 GB laptop,…

22:00
2026-08-23
fratepietro.com
artificial-intelligence

Ferrox v0.9.1: a Rust GGUF engine, measured against llama.cpp

Ferrox v0.9.1, a pure-Rust GGUF inference engine, shipped today with MoE prefill on Apple Metal 2.4x faster, closing the gap to llama.cpp from 2.62x behind to 1.11x on OLMoE-1B-7B. The update also imp…

22:00
2026-08-23
fratepietro.com
artificial-intelligence

Ferrox on Metal: at parity with llama.cpp, and past it

Ferrox, a pure-Rust GGUF inference engine, now runs mixture-of-experts (MoE) prefill on Apple Metal 2.4x faster, reaching 1402 tok/s on OLMoE-1B-7B and closing the gap to llama.cpp from 2.62x behind t…

09:00
2026-08-05
fratepietro.com
artificial-intelligence

Building a Rust Inference Engine That Matches Llama.cpp

Developer Antonello F. released Ferrox, a pure-Rust inference engine that runs open LLMs locally on CPU, Apple Metal, or CUDA, achieving performance parity with llama.cpp on an Apple M2 Pro: 26.9 tok/…

// co-occurs with top 8 entities
// topics top 6 topics