Ramabana – Rama's Arrow
Ramabana, a new coding agent from developer Vedic Reader, is now available via pip, featuring a host that coordinates policy, tools, memory, and routing, with models loaded through the rishi tool supp…
Ramabana, a new coding agent from developer Vedic Reader, is now available via pip, featuring a host that coordinates policy, tools, memory, and routing, with models loaded through the rishi tool supp…
Google engineers built Gemma Translator, an open-source project that runs a fully offline multilingual voice translator on a Raspberry Pi 5 using Google Gemma 4 E2B (2.3B effective / ~5.1B total param…
Google AI Edge's LiteRT and LiteRT-LM enable deployment of Gemma 4 E2B on Raspberry Pi 5, achieving 99 tokens/sec prefill and 9 tokens/sec decode with a peak memory footprint of 1432 MB, powering the …
Garden Reads, an offline e-reader with an on-device large language model powered by Google's Gemma via LiteRT, launches with no account, no tracking, and no subscription. The app offers reading, audio…
Google released LiteRT.js on July 9, 2026, a JavaScript binding of LiteRT that runs .tflite models in browsers via WebAssembly, XNNPACK on CPU, ML Drift over WebGPU, and experimental WebNN for NPUs. G…
A new open-source benchmark, "apple-silicon-llm-bench," reveals that Google's LiteRT-LM runtime outperforms MLX-Swift on the iPhone 17 Pro for Gemma 4 E2B inference, achieving 55.4 tok/s with 4.5x les…
Google launched Coral, a full-stack platform for edge AI that combines an AI-first hardware architecture with a unified developer experience to enable efficient, local AI processing. The platform prov…
The article describes the development of Remora, a privacy-focused dream journaling app that uses Google's Gemma 4 E2B model for on-device AI analysis, avoiding cloud uploads. The engineering team fac…
SafeSMS is a privacy-focused Android application that uses the on-device Gemma 4 E4B AI model to detect SMS-based scams, phishing, and spam without sending data to the cloud. The app analyzes incoming…
LiteRT, a cross-platform framework for on-device AI, enables developers to leverage Neural Processing Units (NPUs) for faster and more efficient AI features like real-time video effects and speech rec…
Arm's Scalable Matrix Extension 2 (SME2) integrates a matrix-compute unit into the CPU, enabling up to 5x faster inference for generative AI workloads on mobile devices. It highlights how Google AI Ed…
LiteRT-LM, part of Google AI Edge, is an optimized runtime engine for deploying Gemma 4 models on-device across platforms like Chrome, ChromeOS, and Pixel Watch. It achieves high performance through a…
The Google Tensor ML SDK has moved from an Experimental Access Program to Beta, now integrating with LiteRT to provide a unified API for deploying machine learning models on the Tensor Processing Unit…