Perplexity Will Open Source Its Faster Lily AI Engine For Apple Silicon Perplexity announced it will open source Lily, a local AI engine built for Apple silicon and the Qwen3.6-35B-A3B model, which uses a Rust runtime and custom Metal kernels without PyTorch or MLX. Perplexity reports Lily averaged 23 percent faster prompt processing and 35 percent faster token generation than MLX-LM on an M5 Max MacBook Pro with 128GB of unified memory, though the code is not yet available and claims rely on internal testing. BrianFagioli writes: Perplexity has built a local artificial intelligence engine designed specifically for Apple silicon and the Qwen3.6-35B-A3B model. Called Lily, the engine uses a Rust runtime and custom Metal kernels, with neither PyTorch nor MLX in its execution path. Perplexity says Lily averaged 23 percent faster prompt processing and 35 percent faster token generation than MLX-LM on an M5 Max MacBook Pro with 128GB of unified memory. Lily is more specialized than MLX-LM, which supports a much wider range of models and architectures. Perplexity says it plans to release Lily as open source, but the code is not available yet, leaving its performance claims dependent on internal testing for now. Read more of this story https://entertainment.slashdot.org/story/26/09/02/2235237/perplexity-will-open-source-its-faster-lily-ai-engine-for-apple-silicon?utm source=rss1.0moreanon&utm medium=feed at Slashdot.