ML Drift: Next-Gen GPU AI/ML Inference at the Edge Google's AI Edge Team open-sourced ML Drift, a cross-platform on-device GPU compute engine for AI/ML inference, under the Apache 2.0 license. ML Drift abstracts OpenGL ES, OpenCL, Metal, and WebGPU hardware APIs, serves as the core GPU acceleration engine inside LiteRT, and adds tensor virtualization for unified shaders, an extensible custom op framework with an agentic SKILL.md guide, and 5D tensor support that the legacy TFLite GPU delegate lacked. The Google AI Edge Team is excited to announce the open-source release of ML Drift https://github.com/google-ai-edge/ml-drift , our high-performance, cross-platform, on-device GPU compute engine specifically built for on-device AI/ML inference, under the Apache 2.0 license. By abstracting hardware and low-level API complexities of on-device GPUs across OpenGL ES, OpenCL, Metal, and WebGPU, ML Drift empowers developers to build real-time, interactive ML experiences from advanced video effects to generative AI across multiple platforms. Serving as the core GPU acceleration engine within LiteRT https://developers.google.com/edge/litert , ML Drift is also available as a standalone library for custom graphics and inference runtimes providing a unified foundation to deliver peak performance everywhere. Unlike datacenter inference, where models run on homogeneous, predictable accelerator clusters, deploying GPU-accelerated AI to edge devices is defined by significant hardware diversity. Developers are faced with a wide range of GPU architectures, driver versions, and low-level APIs, not knowing a priori on which specific hardware the application will run on. The TensorFlow Lite GPU delegate established foundational GPU acceleration, but the ecosystem has since evolved: current on-device workloads now encompass a broad spectrum, ranging from real-time computer vision, audio, and depth processing models to high-parameter generative AI. These architectures push consumer silicon to its limits, creating compute and memory bottlenecks that legacy runtimes weren't designed to solve. This shift necessitates a universal engineering foundation of ensuring portability and performance across both classical and next-generation models. Drawing on optimization principles discussed in "Speed is all you need" https://research.google/blog/speed-is-all-you-need-on-device-acceleration-of-large-diffusion-models-via-gpu-aware-optimizations/ , we've built a single, unified framework that delivers stability for traditional ML and peak performance for the most advanced generative AI, ensuring that your models are running on our most efficient, forward-looking GPU compute engine. To achieve both peak performance and expanded model coverage across classical and generative architectures, ML Drift introduces core architectural changes and structural upgrades over the legacy TFLite GPU delegate. Unified Shaders via Tensor Virtualization: Historically, maintaining optimized shaders for the TFLite GPU delegate required hardcoding logical tensor mappings directly to physical GPU objects such as textures and buffers separately across OpenGL, OpenCL, and Metal backends. ML Drift introduces tensor virtualization, a core architectural paradigm that decouples a tensor's logical representation from its physical allocation on the GPU. By having dynamic shader templates resolve and translate coordinates during the compiler’s initialization phase, this unified shader model eliminates the need to maintain disparate, backend-specific shader codebases, adding minimal runtime overhead while preserving cross-platform model portability. Extensible Custom Op Framework: ML Drift adopts a modernized custom op framework that provides direct registration APIs with low-level shading language access. To accelerate this, ML Drift includes an agentic SKILL.md guide, enabling coding agents to author, register, and verify performant custom shaders in minutes. This provides developers with granular control to integrate specialized model blocks directly into the execution graph, drastically reducing the barrier to deploying proprietary architectures. 5D Tensor Support: The TFLite GPU delegate is structurally hardcoded to 4D tensors, forcing developers to use layout hacks for complex models requiring 5D tensors. ML Drift reflected this long overdue functionality in the new LiteRT ML Drift GPU accelerator https://developers.google.com/edge/litert/next/gpu to enable 5D tensor support out-of-the-box. With this, LiteRT executes complex workloads like 3D convolutional networks for volumetric/spatial AI and spatiotemporal models e.g., YOLO 11n, MobileViT v2, and Swin Transformer v2 directly on edge GPUs. Performance Upgrades for Classical Models: As the direct successor to our legacy GPU backend, ML Drift provides a strict performance upgrade for existing classical workloads. By decoupling the runtime from hardware-specific constraints and improving kernel execution efficiency, models that were already running on TFLite GPU gain immediate performance improvements. Migration is designed to be straightforward, with legacy workloads maintaining or exceeding previous benchmarks. Stage-aware Optimizations for Edge LLMs: Autoregressive LLMs have two different computational workloads during inference: the compute-bound KV cache prefill stage and the memory-bandwidth-bound decode stage, generating tokens one by one. ML Drift dynamically switches kernels and layout configurations depending on the active execution phase. During the decoding stage, it utilizes a custom, convolution-aligned KV cache layout and applies aggressive in-kernel activation quantization to bypass redundant memory roundtrips. Broadening across Platforms Desktop Previews : In today’s emerging agentic coding era, highly responsive desktop inference is becoming vital for powering local coding agents and workstation workflows. While we originally engineered ML Drift’s WebGPU backend for in-browser acceleration, Dawn https://dawn.googlesource.com/dawn Chromium's WebGPU implementation allowed us to compile that exact same codebase natively outside the browser. This unified API model enables ML Drift to run on Windows and Linux, leveraging WebGPU's modern, native hardware abstraction to bypass DirectX and Vulkan fragmentation, while complementing our existing high-performance native Metal backend on macOS. Our primary focus remains delivering lightweight, production-grade inference where resource constraints are tightest: mobile and edge devices. However, local development requires flexibility across developer machines. Below is an early snapshot of ML Drift running Gemma models on workstation hardware, illustrating how our unified runtime scales across environments. Across edge and workstation environments alike, peak throughput only tells half the story: Memory footprint dictates what workloads can actually run concurrently. In our Gemma benchmarks, ML Drift achieves up to 12% lower memory overhead than other frameworks, freeing up memory for other local tasks. ML Drift is already running in production across millions of devices daily, powering key features across Google's ecosystem, including Chrome https://www.google.com/chrome/ai-innovations/ , YouTube Shorts, Photos https://www.google.com/intl/en us/photos/editing/ , Meet https://workspace.google.com/blog/productivity-collaboration/bringing-your-best-self-more-meaningful-connections-google-meet , and the recently launched AI Edge Gallery https://developers.google.com/edge/gallery . YouTube Shorts migrated its segmentation-based effects to use ML Drift. This delivered up to a 40% reduction in average frame latency across both Android and iOS, enabling creators to apply high-fidelity visual effects in real time without dropping camera frames. To deliver fast, seamless on-device photo editing, the Google Photos team integrated ML Drift across its computational photography and segmentation pipelines. This delivered up to 2 sec speedup compared to the legacy GPU delegate, making advanced photo enhancements feel instantaneous. Google Chrome integrated ML Drift into its AI runtime infrastructure to power native hardware-accelerated Gemini Nano models for Built-In AI https://developer.chrome.com/docs/ai/built-in APIs including the Prompt, Summarize, and Writer APIs . Across desktop platforms, ML Drift enables high-efficiency on-device AI across a growing ecosystem of web applications from e-commerce review summarization to enterprise customer workflows: Beyond Google products, leading developer partners have adopted ML Drift to bring desktop grade capabilities to mobile users: Adobe Lightroom https://lightroom.app.link/PAkdc2Jzx6b and Adobe Photoshop https://adobephotoshop.app.link/YSkuYkUzx6b upgraded key AI features, including Select Subject, Select Sky, and Adaptive Portrait, with ML Drift, achieving up to 30% faster on-device performance for pro-quality photo editing on mobile. Snap integrated LiteRT to power on-device inference on GPU across Android devices. ML Drift GPU accelerator brought 30% model latency improvements for ML-powered face and style generator effects in Snapchat lenses. ML Drift not only improves user experiences based on existing models, but also expands capabilities to run even large diffusion models. Maximizing edge AI efficiency requires deep alignment between software architectures and physical silicon. Throughout ML Drift's development, Google worked closely with silicon and IP partners to enable optimizations that maximize hardware utilization across modern GPUs: By working hand-in-hand with silicon leaders, ML Drift ensures that state-of-the-art model optimizations translate into battery-conscious, high-frame rate performance on user devices. With the launch of ML Drift, the legacy TFLite GPU delegate will no longer receive new feature updates. We encourage all developers to transition to the LiteRT ML Drift GPU accelerator which offers full backwards compatibility for your existing models while immediately unlocking modern performance gains across Android, iOS, Web, and desktop. For Android developers using unbundled runtimes to minimize app binary size, ML Drift acceleration is available today in standalone LiteRT packages and is coming soon to LiteRT in Google Play Services https://developers.google.com/edge/litert/android . Whether you are integrating on-device models into an app or authoring custom GPU shaders, ML Drift provides two straightforward paths: Community feedback and Contributions ML Drift is fully open source under the Apache 2.0 license. We welcome contributions, RFCs, and community feedback: Acknowledgments: