11:26
2026-08-26
github.com
artificial-intelligence
Axera AX8850 LLM running ggufs
A custom llama.cpp backend (ggml-axcl) now runs Qwen3-0.6B directly from GGUF files on an Axera AX8850 NPU accelerator card (M5Stack LLM-8850) hosted on a Raspberry Pi 5, achieving 1.3-2.7 tokens per …