cd /news/ai-agents/sub-15ms-local-decisions-running-lay… · home › topics › ai-agents › article
[ARTICLE · art-148568] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

Sub-15ms Local Decisions: Running Laya Models with Native MLX on Apple Silicon

A developer released laya-mlx, a native Apple MLX runtime for Laya typed decision models that performs structured routing decisions in 7–14 ms on an M3 Max. The project strips out autoregressive text generation and PyTorch dependencies to give local agent workflows deterministic, on-device decision-making, targeting engineers who need sub-20ms routing for autonomous agents, real-time code assistants, and edge devices.

by read1 min views1 publishedOct 10, 2026

Stop burning tokens and introducing multi-second cloud latency just to make structured routing decisions. laya-mlx is a native Apple MLX runtime for Laya typed decision models that clocks in at an astonishing 7–14 ms on an M3 Max. By stripping away autoregressive text generation and heavy PyTorch dependencies, it gives local AI workflows instant, deterministic decision-making entirely on-device.

Get up and running with laya-mlx in just a few lines of code:

pip install mlx laya-mlx
python
from laya_mlx import LayaDecisionModel

model = LayaDecisionModel.from_pretrained("mizore/laya-decision-v1")

context = "User requested database query optimization on production cluster."
decision = model.decide(
    context=context,
    schema=["escalate_to_dba", "run_auto_explain", "reject_request"]
)

print(f"Decision: {decision.action} (Latency: {decision.latency_ms:.2f}ms)")

Most modern agentic architectures overuse massive 7B+ LLMs for tasks that are fundamentally multi-class classifications or rigid tool routers. Invoking an LLM via cloud API adds 500–1500 ms of latency and burns cash; running a 7B model locally via Ollama or vLLM consumes gigabytes of VRAM and still takes hundreds of milliseconds.

laya-mlx is built for engineers building autonomous local agent systems, real-time code assistants, and edge devices where sub-20ms routing is mandatory. If you need reliable, typed branch logic without spinning up a heavy generative pipeline, this is the architecture to watch.

Are you optimizing your local AI infrastructure for speed and efficiency? Star the project on GitHub: mizorewww/laya-mlx.

Follow 'Local AI & Infra Daily' for daily deep dives into the fastest runtimes, local model optimizations, and open-source AI infrastructure!

── more in #ai-agents 4 stories · sorted by recency
── more on @laya-mlx 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/sub-15ms-local-decis…] indexed:0 read:1min 2026-10-10 · —