Sub-15ms Local Decisions: Running Laya Models with Native MLX on Apple Silicon A developer released laya-mlx, a native Apple MLX runtime for Laya typed decision models that performs structured routing decisions in 7–14 ms on an M3 Max. The project strips out autoregressive text generation and PyTorch dependencies to give local agent workflows deterministic, on-device decision-making, targeting engineers who need sub-20ms routing for autonomous agents, real-time code assistants, and edge devices. Stop burning tokens and introducing multi-second cloud latency just to make structured routing decisions. laya-mlx is a native Apple MLX runtime for Laya typed decision models that clocks in at an astonishing 7–14 ms on an M3 Max. By stripping away autoregressive text generation and heavy PyTorch dependencies, it gives local AI workflows instant, deterministic decision-making entirely on-device. Get up and running with laya-mlx in just a few lines of code: pip install mlx laya-mlx python from laya mlx import LayaDecisionModel Load a pre-trained typed decision model natively into Apple Silicon unified memory model = LayaDecisionModel.from pretrained "mizore/laya-decision-v1" Provide input context and define the decision target context = "User requested database query optimization on production cluster." decision = model.decide context=context, schema= "escalate to dba", "run auto explain", "reject request" print f"Decision: {decision.action} Latency: {decision.latency ms:.2f}ms " Output: Decision: run auto explain Latency: 9.42ms Most modern agentic architectures overuse massive 7B+ LLMs for tasks that are fundamentally multi-class classifications or rigid tool routers. Invoking an LLM via cloud API adds 500–1500 ms of latency and burns cash; running a 7B model locally via Ollama or vLLM consumes gigabytes of VRAM and still takes hundreds of milliseconds. laya-mlx is built for engineers building autonomous local agent systems, real-time code assistants, and edge devices where sub-20ms routing is mandatory. If you need reliable, typed branch logic without spinning up a heavy generative pipeline, this is the architecture to watch. Are you optimizing your local AI infrastructure for speed and efficiency? Star the project on GitHub: mizorewww/laya-mlx https://github.com/mizorewww/laya-mlx . Follow 'Local AI & Infra Daily' for daily deep dives into the fastest runtimes, local model optimizations, and open-source AI infrastructure