cd/entity/MLX· home entities MLX
grep -l @mlx /news/*.json | wc -l → 114

MLX

mentions 114 type Organization page 5/6 feed RSS

// recent coverage 114 mentions

01:32
2026-07-10
digital-foundry-eight.vercel.app
large-language-models

I benchmarked every model that fits on an iPhone

An independent benchmark of on-device LLMs on iPhone A17 Pro found Apple's system model achieves ~149 tok/s with only 12MB peak app memory, while 4B-class open models like Qwen3 4B and Llama 3.2 3B tr…

03:30
2026-07-08
byteiota.com
large-language-models

Ollama 0.31: Gemma 4 Runs 90% Faster on Apple Silicon

Ollama 0.31 introduces multi-token prediction for Gemma 4 on Apple Silicon, cutting token generation time by 90% on Aider's coding benchmark. The update automatically accelerates local coding agents w…

19:15
2026-07-07
github.com
large-language-models

Local AI is re-reading its own prompt

A local AI assistant called Kira discovered that nearly half of its model time was spent re-reading boilerplate prompt text—a 'prefill tax'—after migrating from Ollama to Apple's MLX framework, which …

06:59
2026-07-04
maloyan.xyz
large-language-models

Running Qwen 3.6 Locally on a Mac Mini M4 with 16GB RAM

Qwen open-sourced the 35-billion parameter Mixture of Experts model Qwen 3.6-35B-A3B, which activates only 3 billion parameters per token and runs on a $599 Mac Mini M4 with 16GB RAM at 17 tok/s with …

23:26
2026-07-03
dev.to
developer-tools

Run Claude Code locally for free: mlx-serve on Apple Silicon

A developer released mlx-serve, a native Zig server for MLX-format language models on Apple Silicon, enabling local, free, and private use of AI coding assistants like Claude Code. The server exposes …

08:23
2026-06-30
xda-developers.com
large-language-models

Good article about local LLM on MacBook Air

Ollama's new MLX engine enables local LLM inference on MacBook Air at twice the speed, making powerful AI accessible on consumer hardware without cloud dependency.…

19:00
2026-06-29
bart.degoe.de
artificial-intelligence

Running a coding agent locally for when the budget runs out

Uber capped employee AI coding tool spending at $1,500 per month in June 2026 after burning through its entire 2026 budget in four months, with Claude Code adoption surging from a third of engineers t…

11:32
2026-06-27
akshit.org
large-language-models

William: A tiny poetry model in the browser

William, a tiny poetry language model trained by a developer, runs entirely in the browser using ONNX Runtime Web. The 6-layer transformer was trained on the Gutenberg Poetry Corpus and fine-tuned on …

00:00
2026-06-26
squish.run
large-language-models

I Couldn't Find a Local LLM Tool Fast Enough, So I Built My Own

A developer built Squish, a local LLM inference server for Apple Silicon, after finding existing tools like Ollama too slow for generating git commit messages. Squish achieves up to 9.8× faster perfor…

08:22
2026-06-22
github.com
developer-tools

Show HN: Gingerpaw : A voice dictation and agent workspace app

Gingerpaw, a new macOS app, launches a multi-agent coding workspace with on-device voice dictation and a talking ginger cat notification system. The app runs multiple coding agents like Claude Code an…

11:47
2026-06-21
vettedconsumer.com
large-language-models

Show HN: Local LLM Hardware Calculator

A new Local LLM Hardware Calculator helps users estimate memory requirements for running large language models on their own hardware, factoring in weights, KV cache, and overhead. The tool also compar…

20:58
2026-06-20
vettedconsumer.com
large-language-models

Qwen3-30B-A3B: The Open Model Most People Should Actually Run

Alibaba's Qwen team released Qwen3-30B-A3B, a Mixture-of-Experts model with 30.5 billion total parameters but only 3.3 billion active per token, enabling it to run on a single 24 GB graphics card at s…

20:18
2026-06-18
console.cocore.dev
ai-infrastructure

Co/Core: An AI Cooperative

Co/Core launches an AI cooperative where members share their own compute hardware to run AI inference jobs, using an open standard and public receipts to ensure transparency. The platform supports exi…

21:00
2026-06-17
veso.ai
machine-learning

What We Learned Validating JEPA on a Laptop

Researchers reproduced the core mechanisms of two 2025-26 JEPA papers (CrossJEPA and LeJEPA) at toy scale on an M5 MacBook Air, confirming that the self-supervised representation learning methods work…

16:42
2026-06-16
flox.dev
machine-learning

Training NanoGPT on Slurm with a Nix-Pinned Environment

A researcher using nanoGPT on a MacBook Pro faces dependency failures when moving to a GPU cluster, prompting a solution using Nix and Flox to create reproducible, cross-platform runtime environments …

← prev page 5 / 6 next →
// co-occurs with top 8 entities
// topics top 6 topics