cd/sources/fratepietro-auto-discovered· home› sources› Fratepietro (auto-discovered)
cat /sources/fratepietro-auto-discovered.feed | wc -l → 11

Fratepietro (auto-discovered)

articles 11 domain fratepietro.com → feed RSS
22:00
2026-09-27
fratepietro.com
artificial-intelligence

Frink Inference: the Rust alternative to llama.cpp

Frink, a pure-Rust inference engine for GGUF models created by developer Antonello F, now runs 100 architectures on its generic path with evidence, each verified against llama.cpp's own logits via lib…

22:00
2026-09-22
fratepietro.com
large-language-models

LLM Inference: a learning roadmap

A new learning roadmap breaks LLM inference into five layers — the call itself, transformer computation, hardware limits, optimizations, and serving engines — explaining that prefill builds the KV cac…

22:00
2026-09-17
fratepietro.com
large-language-models

A 27B model in 5.5 GB: Bonsai ternary on Ferrox, locally

Ferrox v0.23.0, released today, adds support for PrismML's Ternary-Bonsai-2-27B, a 27-billion-parameter model quantized to 1.75 bits per weight that occupies 5.5 GB on disk and runs on a 16 GB laptop,…

22:00
2026-09-02
fratepietro.com
developer-tools

gitgui: a real git GUI inside cmux, next to Pi

Developer Antonello F. released gitgui, a Rust binary that renders a full git GUI—commit graph, staged file list, inline diffs, and hunk buttons—inside terminal panes using the kitty graphics protocol…

22:00
2026-08-23
fratepietro.com
ai-infrastructure

Frink on Metal: at parity with llama.cpp, and past it

Frink, a pure-Rust GGUF inference engine written by Antonello Frink, reached parity with llama.cpp on Apple Metal decode and passed it on prefill at version v0.12.0, with OLMoE-1B-7B prefill rising fr…

22:00
2026-08-23
fratepietro.com
artificial-intelligence

Ferrox v0.9.1: a Rust GGUF engine, measured against llama.cpp

Ferrox v0.9.1, a pure-Rust GGUF inference engine, shipped today with MoE prefill on Apple Metal 2.4x faster, closing the gap to llama.cpp from 2.62x behind to 1.11x on OLMoE-1B-7B. The update also imp…

22:00
2026-08-23
fratepietro.com
artificial-intelligence

Ferrox on Metal: at parity with llama.cpp, and past it

Ferrox, a pure-Rust GGUF inference engine, now runs mixture-of-experts (MoE) prefill on Apple Metal 2.4x faster, reaching 1402 tok/s on OLMoE-1B-7B and closing the gap to llama.cpp from 2.62x behind t…

09:00
2026-08-05
fratepietro.com
artificial-intelligence

Building a Rust Inference Engine That Matches Llama.cpp

Developer Antonello F. released Ferrox, a pure-Rust inference engine that runs open LLMs locally on CPU, Apple Metal, or CUDA, achieving performance parity with llama.cpp on an Apple M2 Pro: 26.9 tok/…

22:00
2026-07-21
fratepietro.com
artificial-intelligence

Running GLM-5.2 Locally with Rondine and Pi

Rondine, a tool that detects hardware and configures inference engines, enables running GLM-5.2, a 744B-parameter Mixture-of-Experts coding model requiring approximately 245GB of memory, locally on a …