Long time local model fan here. I’ve been building Anvil, an iPhone app that runs everything on the phone, and wanted to share how it fits.
Chat is Gemma 4 E2B through LiteRT LM. Photo and document questions use the same model, so it can read and translate a sign or a menu with no signal.
Images come from an SDXL model converted to Core ML, running on the Neural Engine at 1024x1024, as many as you want, even in airplane mode. The grid above was all made on the phone.
The memory trick is boring but it works: the two models are never loaded at the same time. That’s also why it needs an iPhone 15 Pro or newer.
No account, and nobody else sees what you ask. Web search and iCloud backup are optional and are the only parts that touch the network.
Version 4.1 is on public TestFlight and testers get Anvil Pro free: Join the Anvil – Offline AI Assistant beta - TestFlight - Apple
If you’ve pushed other diffusion models onto iOS, I’d love to hear what compression settings kept quality up.