{"slug": "sub-15ms-local-decisions-running-laya-models-with-native-mlx-on-apple-silicon", "title": "Sub-15ms Local Decisions: Running Laya Models with Native MLX on Apple Silicon", "summary": "A developer released laya-mlx, a native Apple MLX runtime for Laya typed decision models that performs structured routing decisions in 7–14 ms on an M3 Max. The project strips out autoregressive text generation and PyTorch dependencies to give local agent workflows deterministic, on-device decision-making, targeting engineers who need sub-20ms routing for autonomous agents, real-time code assistants, and edge devices.", "body_md": "Stop burning tokens and introducing multi-second cloud latency just to make structured routing decisions. **laya-mlx** is a native Apple MLX runtime for Laya typed decision models that clocks in at an astonishing 7–14 ms on an M3 Max. By stripping away autoregressive text generation and heavy PyTorch dependencies, it gives local AI workflows instant, deterministic decision-making entirely on-device.\n\nGet up and running with `laya-mlx` in just a few lines of code:\n\n```\npip install mlx laya-mlx\npython\nfrom laya_mlx import LayaDecisionModel\n\n# Load a pre-trained typed decision model natively into Apple Silicon unified memory\nmodel = LayaDecisionModel.from_pretrained(\"mizore/laya-decision-v1\")\n\n# Provide input context and define the decision target\ncontext = \"User requested database query optimization on production cluster.\"\ndecision = model.decide(\n    context=context,\n    schema=[\"escalate_to_dba\", \"run_auto_explain\", \"reject_request\"]\n)\n\nprint(f\"Decision: {decision.action} (Latency: {decision.latency_ms:.2f}ms)\")\n# Output: Decision: run_auto_explain (Latency: 9.42ms)\n```\n\nMost modern agentic architectures overuse massive 7B+ LLMs for tasks that are fundamentally multi-class classifications or rigid tool routers. Invoking an LLM via cloud API adds 500–1500 ms of latency and burns cash; running a 7B model locally via Ollama or vLLM consumes gigabytes of VRAM and still takes hundreds of milliseconds.\n\n`laya-mlx` is built for engineers building autonomous local agent systems, real-time code assistants, and edge devices where sub-20ms routing is mandatory. If you need reliable, typed branch logic without spinning up a heavy generative pipeline, this is the architecture to watch.\n\nAre you optimizing your local AI infrastructure for speed and efficiency? Star the project on GitHub: [mizorewww/laya-mlx](https://github.com/mizorewww/laya-mlx).\n\n**Follow 'Local AI & Infra Daily'** for daily deep dives into the fastest runtimes, local model optimizations, and open-source AI infrastructure!", "url": "https://wpnews.pro/news/sub-15ms-local-decisions-running-laya-models-with-native-mlx-on-apple-silicon", "canonical_source": "https://dev.to/hui_feng_f2247629b1d2be00/sub-15ms-local-decisions-running-laya-models-with-native-mlx-on-apple-silicon-15bk", "published_at": "2026-10-10 01:52:23+00:00", "updated_at": "2026-10-10 01:59:13.150785+00:00", "lang": "en", "topics": ["ai-agents", "ai-infrastructure", "ai-tools", "machine-learning", "developer-tools"], "entities": ["laya-mlx", "Apple MLX", "Laya", "M3 Max", "GitHub", "Ollama", "vLLM", "PyTorch"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/sub-15ms-local-decisions-running-laya-models-with-native-mlx-on-apple-silicon", "markdown": "https://wpnews.pro/news/sub-15ms-local-decisions-running-laya-models-with-native-mlx-on-apple-silicon.md", "text": "https://wpnews.pro/news/sub-15ms-local-decisions-running-laya-models-with-native-mlx-on-apple-silicon.txt", "jsonld": "https://wpnews.pro/news/sub-15ms-local-decisions-running-laya-models-with-native-mlx-on-apple-silicon.jsonld"}}