cd/entity/Apple Silicon· home entities Apple Silicon
grep -l @apple silicon /news/*.json | wc -l → 150

Apple Silicon

mentions 150 type Organization page 5/8 feed RSS

// recent coverage 150 mentions

00:00
2026-07-08
tomtunguz.com
artificial-intelligence

The AI Preflight Check

A developer built an AI agent with a memory architecture that runs preflight instructions, retrieving relevant skills from a library of ~90 workflow files before executing tasks. The system uses a loc…

00:00
2026-07-08
runagentrun.co.uk
artificial-intelligence

mistral.rs v0.9.0 outpaces llama.cpp on CPU

Mistral.rs released version 0.9.0 on 7 July, claiming up to 1.8× faster CPU decoding than llama.cpp on both x86 and ARM hardware, challenging the de-facto standard for local LLM inference. The speedup…

11:52
2026-07-07
github.com
machine-learning

Show HN: TurboQuant for mlx-lm (Apple Silicon)

A developer released TurboQuant, a pip-installable adapter for mlx-lm on Apple Silicon that uses a randomized Hadamard transform for data-oblivious, calibration-free quantization of weights and KV cac…

23:26
2026-07-03
dev.to
developer-tools

Run Claude Code locally for free: mlx-serve on Apple Silicon

A developer released mlx-serve, a native Zig server for MLX-format language models on Apple Silicon, enabling local, free, and private use of AI coding assistants like Claude Code. The server exposes …

12:30
2026-07-01
basecompute.co
ai-infrastructure

BaseRT, A fast inference runtime for local AI on Apple Silicon

BaseCompute released BaseRT, a fast inference runtime for local AI on Apple Silicon, claiming up to 35% faster decode and 78% faster prefill on an Apple M4 Pro with 4-bit quantization. The runtime all…

00:22
2026-07-01
ollama.com
large-language-models

Faster Gemma 4 on MLX with multi-token prediction

Gemma 4 in Ollama 0.31 generates tokens nearly 90% faster on Apple Silicon using multi-token prediction (MTP), which employs a small draft model to propose multiple tokens that the main model verifies…

16:32
2026-06-30
github.com
ai-agents

Show HN: Agent Sandbox Options

A collection of open-source sandbox tools for AI coding agents has been released, offering hardware-isolated microVMs, containers, and isolation harnesses that boot in under 200ms and prevent secret l…

05:32
2026-06-29
squish.run
large-language-models

Squish – The fastest way to run local LLMs on Apple Silicon

Squish, a new local AI agent runtime for Apple Silicon, claims to load models 54× faster than standard paths and serve them faster than Ollama, with full offline capability and no cloud dependencies. …

08:55
2026-06-28
forum.modular.com
artificial-intelligence

You can now run Max AI models on Apple Silicon

Modular announced that MAX models can now run on Apple Silicon GPUs with the 26.4 release, supporting M1 through M5 chips for text LLMs, vision models, and image diffusion models. The company is worki…

11:32
2026-06-27
akshit.org
large-language-models

William: A tiny poetry model in the browser

William, a tiny poetry language model trained by a developer, runs entirely in the browser using ONNX Runtime Web. The 6-layer transformer was trained on the Gutenberg Poetry Corpus and fine-tuned on …

00:00
2026-06-26
squish.run
large-language-models

I Couldn't Find a Local LLM Tool Fast Enough, So I Built My Own

A developer built Squish, a local LLM inference server for Apple Silicon, after finding existing tools like Ollama too slow for generating git commit messages. Squish achieves up to 9.8× faster perfor…

← prev page 5 / 8 next →
// co-occurs with top 8 entities
// topics top 6 topics