How to Add Offline Speech, Image Understanding, and Image Generation to Your App in 2026 A developer has released OGAD (Off Grid AI Desktop), a local gateway that exposes offline vision, transcription, image-generation and text-to-speech capabilities to apps through a single HTTP address on port 7878. The tool routes each task to a separate locally downloaded model — vision-capable chat models for image understanding, transcription models for audio, image-generation models for illustrations and a supported speech runtime for spoken output — so a configured workflow can run without an online AI provider. The guide notes the verified Windows stable package lacks the same local speech-output runtime, so a Mac is required for the full capability set. A small app can do more than send text to a chatbot. It can explain a screenshot, turn a recording into text or create an illustration while the required models run on the user's computer. OGAD Off Grid AI Desktop exposes those local capabilities through one gateway. Your app uses a separate model route for each job, with a shared HTTP address. Download the models and required assets first; then a configured local workflow can run without an online AI provider. Download OGAD for Mac or Windows https://getoffgridai.co/desktop/ | Reader or user task | API route | Local model needed | |---|---|---| | Explain an image | POST /v1/chat/completions with image content | Vision-capable chat model | | Transcribe a recording | POST /v1/audio/transcriptions | Transcription model | | Create an illustration | POST /v1/images/generations | Image-generation model | | Read text aloud | POST /v1/audio/speech | Supported speech-output runtime and voice | These are separate capabilities. A text-only model does not become a vision model because the request contains an image. Use a Mac for the complete set in this guide. The verified Windows stable package does not include the same local speech-output runtime. Windows can use supported local text/vision, transcription and image-generation routes, but check the installed model and runtime for each task. Core inference and the gateway do not require Pro capture. Start with one capability before building a UI around several at once. In OGAD, download and select a supported local vision model in Models . Open Gateway and check the local address. The example below uses the normal 7878 port; replace it if your app shows another one. Save a small PNG image as example.png . This Python 3 script sends the local image as a data URL, so it does not need a remotely hosted image: python import base64 import json from pathlib import Path from urllib.request import Request, urlopen BASE = "http://127.0.0.1:7878" def request json path, payload=None : data = None if payload is None else json.dumps payload .encode "utf-8" request = Request BASE + path, data=data, headers={"Content-Type": "application/json"} with urlopen request, timeout=240 as response: return json.load response models = request json "/v1/models" "data" model = next item for item in models if item.get "kind" == "vision" and not item.get "remote" , None if model is None: raise SystemExit "Select a downloaded local vision model in OGAD first." encoded = base64.b64encode Path "example.png" .read bytes .decode "ascii" result = request json "/v1/chat/completions", { "model": model "id" , "messages": {"role": "user", "content": {"type": "text", "text": "Describe the visible image. Mark unclear details."}, {"type": "image url", "image url": {"url": "data:image/png;base64," + encoded}}, } , "max tokens": 160, "stream": False, } print result "choices" 0 "message" "content" Compare the description with the source image. Image understanding can make mistakes, especially with small text or ambiguous details. Your app should let the user check the image beside the answer. For image generation, select a downloaded image model and send a request such as: { "prompt": "A simple watercolor illustration of a quiet reading desk, no text", "size": "512x512", "response format": "b64 json" } Send it to /v1/images/generations . A successful JSON result contains an image in data 0 .b64 json ; decode it before displaying or saving it. Use a size supported by the chosen model. Do not assume a cloud image API's entire parameter set is supported here. For spoken output on a supported Mac, prepare a local voice in OGAD first. The speech endpoint accepts JSON with input text and returns WAV audio by default. Your app must read those bytes as audio rather than trying to parse them as JSON. Transcription uses a multipart file field, not a JSON string containing the recording's path. The running /docs reference gives each route's format. Show useful loading and error states. The first request can include a model load, and a model that does not fit the available memory cannot be fixed by a longer HTTP timeout alone. Keep the first interaction small and avoid firing every model route at once. Use local selections for every capability in an offline workflow. One configured remote provider can make an otherwise local app depend on internet. Complete first-use asset downloads before the offline check. The gateway listens on network interfaces and its inference endpoints do not require an API key. These examples use 127.0.0.1 on the same computer. Keep the host on a trusted network and do not expose this port to the public internet. These API routes are present in OGAD 0.0.51 https://github.com/off-grid-ai/OGAD/releases/tag/v0.0.51 . The running gateway also serves its API reference at /docs . Download OGAD https://getoffgridai.co/desktop/ and connect your app to one image question, one recording or one generated illustration. Add the next capability after the first gives a result the user can inspect.