cd /news/artificial-intelligence/aneforge-program-the-apple-neural-en… · home topics artificial-intelligence article
[ARTICLE · art-76113] src=aneforge.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

ANEForge: Program the Apple Neural Engine Directly, from Python

ANEForge, a new open-source Python package, compiles tensor graphs directly into programs for the Apple Neural Engine (ANE) and runs them without CoreML, enabling both inference and training on the chip Apple designed only for inference. The package achieves a 6.1× speedup over CoreML for a sample workload (0.33 ms vs. 2.0 ms), runs the MLPerf reference ResNet-50 at 76.44% fp16 accuracy, and supports LLMs up to 27B parameters, FFTs, and BLAS operations entirely on the ANE. ANEForge dispatches through the same private aned stack CoreML uses internally, from a normal user process, and cross-compiles for 28 Apple silicon targets from M1 to M5.

read3 min views1 publishedJul 27, 2026
ANEForge: Program the Apple Neural Engine Directly, from Python
Image: source

ANEForge compiles a tensor graph into one program for the Apple Neural Engine and runs it from Python, without CoreML. The same path trains on the engine, optimizer included. The compute unit is fixed, so a model never falls to CPU/GPU.

0.33 ms on the engine

6.1× faster, 2.0 ms

Simulations and learning, run and trained on the engine. #

Not just neural nets. Each of these compiles to ANE programs and runs end to end on the Neural Engine.

Training, LLMs, native layers, compression, and more. #

Apple exposes the Neural Engine only through CoreML, and only for inference. ANEForge dispatches through the same private aned

stack that CoreML uses internally, from a normal user process.

Training runs on the engine CIFAR-10 → 71%

The forward pass, backward pass, and Adam update all compile to ANE programs, so a model trains end to end on the engine.

LLMs on the engine 27B on a pure ANE

Prefill and KV-cache decode for Llama/Qwen, exact speculative decoding, Mixture-of-Experts from GGUF, and the hybrid Qwen3.5-27B (DeltaNet + attention), decoding end to end on the engine.

MLPerf on the engine passes the checker

The MLPerf reference ResNet-50 runs pure-ANE and passes the upstream MLCommons submission_checker

at the reference accuracy (fp16 76.44%, equal to fp32). One command reproduces it.

Layers CoreML can’t reach +19 bridge ops

af.sdpa

drives fused attention directly, what CoreML never emits. Sort, argmax

, topk

, geometry too.

Streaming weight compression 4× smaller

int8, int4-LUT, or sparse weights via the dequant path, accuracy-gated. KV-cache and optimizer state stay resident.

Run arbitrary workloads FFT, BLAS

Not just neural nets. FFTs, linear algebra, and fluid sims compile straight to the engine.

Cross-compile across silicon 28 targets

Lower and gate a graph for any M1 to M5 ANE from one machine, and estimate its latency without running it.

CoreML is the only public route to the engine, and all it decides is whether to use it. #

Path On the ANE No CoreML Trains on it
CoreML / coremltools scheduler chooses no no
| MLX, PyTorch (MPS) | no (GPU) | yes | on the GPU |
| ANEForge | yes (direct) | yes | yes |

Where it lands: #

the reference ResNet-50 against official MLPerf.

The MLPerf reference ResNet-50 on the engine, next to official MLPerf Inference results. Every competitor number is sourced to a public mlcommons/inference_results

submission; the ANE row is unofficial and self-measured.

System SingleStream p90 Offline Power class
Apple Neural Engineunofficial 0.773 ms 1,149/s single-digit W · SoC block
NVIDIA Jetson AGX Orinv3.1 · edge 0.640 ms 6,424/s 15–60 W module
NVIDIA H100, 1 GPUv4.0 · datacenter 88,714/s 350–700 W board

The package, the paper behind it, and a guide to the engine. #

ANEForge

The Python package. Build a graph, compile it to one ANE program, then run or train it. It handles quantized weights, native attention, resident state, and cross-compilation for 28 targets.

[GitHub →](https://github.com/sbryngelson/ANEForge)

[Docs →](https://aneforge.readthedocs.io)

The paper

“Python for direct computation on the Apple Neural Engine.” The preprint behind the numbers on this page, with the dispatch path and the validation against reference in full.

arXiv →

The guide

A reverse-engineering account of the engine: the datapath, the compiler, the program format, the firmware, the kernel driver, and the cross-silicon target tables. Most of it is undocumented by Apple.

Read →

Apple Silicon, macOS 14+, Python 3.10+. #

The e5rt

dispatch shim builds against Apple frameworks on first use. Then browse examples/

, starting with the quickstart.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @aneforge 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/aneforge-program-the…] indexed:0 read:3min 2026-07-27 ·