00:00
2026-08-02
zackproser.com
artificial-intelligence
Practical Advice on Local Inference and Cloud LLMs - August 2026
A developer's two-week experiment with an M5 Max MacBook Pro with 128 GB of unified memory found that DeepSeek V4 Flash weights ran at 6 tokens per second under mainline llama.cpp but 30 to 40 tokens โฆ