cd/entity/MLX· home entities MLX
grep -l @mlx /news/*.json | wc -l → 114

MLX

mentions 114 type Organization page 1/6 feed RSS

// recent coverage 114 mentions

03:09
2026-09-08
byteiota.com
artificial-intelligence

Perplexity Lily: 1.35x Faster Local AI Than MLX on Mac

Perplexity has open-sourced Lily, a Rust and Metal local inference engine that outperforms Apple's MLX framework by 1.35x on decode throughput for Qwen3.6-35B-A3B on Mac, with prefill speeds of 4,156 …

00:00
2026-09-08
unite.ai
artificial-intelligence

How Long Before a Real Crackdown on AI Model Decensoring?

Open-source AI models are increasingly being stripped of their safety filters and redistributed at scale, echoing the warez scene of 1995–2010, according to an analysis on Unite.AI. The practice invol…

13:00
2026-09-06
vettedconsumer.com
large-language-models

Why Is My Local LLM So Slow? The 6 Bottlenecks, in Order

A new guide from Vetted Consumer identifies six bottlenecks that slow local large language models, ranked by impact, with the top cause being the model not fitting in fast memory, which can drop speed…

19:31
2026-09-03
news.ycombinator.com
ai-products

Show HN: Agentray, a macOS agent whose interface is a folder

Agentray, a new macOS app by an unnamed developer, lets users interact with LLMs by dropping files into folders instead of using chat, with the folder name serving as the instruction. The app, current…

02:05
2026-09-03
dev.to
large-language-models

I Tried 4 Models to Save My Self-Improving Agent. All 4 Failed.

A developer testing four language models to improve a self-rewriting AI agent found that none produced promotable edits, despite resolving dozens of bugs. The models, ranging from Qwen3-4B to Mistral …

22:30
2026-09-01
lws.io
artificial-intelligence

My local model setup on an M4 Pro Mac Mini

A developer detailed a local LLM setup on an M4 Pro Mac mini with 48 GB of RAM, running Qwen3.6-35B-A3B-OptiQ-4bit and Gemma-4-E4B-it-OptiQ-4bit via MLX, to avoid cloud API costs, data privacy risks, …

07:00
2026-09-01
haoailab.com
artificial-intelligence

FastH3 on Apple Silicon and DGX Spark

FastH3, MiniMax's video-and-audio generation model, now runs on Apple Silicon via MLX and on NVIDIA's DGX Spark desktop, with two Sparks able to generate one clip together. The release also debuts the…

16:33
2026-08-31
pac.commonsware.com
artificial-intelligence

You Got mlx-serve'd!

Mark Murphy reports that mlx-serve, an Ollama-like inference server with a GUI, delivers a 4x speed increase for running Qwen 3.8 on his 64GB M2 Ultra Mac Studio compared to his previous Ollama setup,…

00:00
2026-08-31
zackproser.com
artificial-intelligence

The gate has to touch the real system

A developer's experiments with an agentic terminal and a fine-tuned model failed because they were measured with self-built instruments that lied, leading to the development of a context engine called…

15:02
2026-08-28
whisprfreely.com
artificial-intelligence

Whispr for Free: How SaaS Is Dead in One Prompt

Whispr Freely, a free macOS dictation app that runs OpenAI's Whisper large-v3-turbo locally on Apple silicon via Apple's MLX framework, is set for public release once Apple's notarization clears, with…

00:00
2026-08-28
mindstudio.ai
ai-infrastructure

Dark Bloom: Rent Out Your Mac for AI Inference and Get Paid

Dark Bloom, a distributed inference network, pays Mac owners to share idle compute for AI inference, with payouts via Stripe and models served through OpenRouter at lower prices. The project, which re…

page 1 / 6 next →
// co-occurs with top 8 entities
// topics top 6 topics