cd /news/ai-tools/how-to-add-offline-speech-image-unde… · home › topics › ai-tools › article
[ARTICLE · art-141632] src=dev.to ↗ pub= topic=ai-tools verified=true sentiment=↑ positive

How to Add Offline Speech, Image Understanding, and Image Generation to Your App in 2026

A developer has released OGAD (Off Grid AI Desktop), a local gateway that exposes offline vision, transcription, image-generation and text-to-speech capabilities to apps through a single HTTP address on port 7878. The tool routes each task to a separate locally downloaded model — vision-capable chat models for image understanding, transcription models for audio, image-generation models for illustrations and a supported speech runtime for spoken output — so a configured workflow can run without an online AI provider. The guide notes the verified Windows stable package lacks the same local speech-output runtime, so a Mac is required for the full capability set.

by read4 min views1 publishedSep 29, 2026

A small app can do more than send text to a chatbot. It can explain a screenshot, turn a recording into text or create an illustration while the required models run on the user's computer.

OGAD (Off Grid AI Desktop) exposes those local capabilities through one gateway. Your app uses a separate model route for each job, with a shared HTTP address. Download the models and required assets first; then a configured local workflow can run without an online AI provider.

Download OGAD for Mac or Windows

Reader or user task API route Local model needed
Explain an image POST /v1/chat/completions with image content Vision-capable chat model
Transcribe a recording POST /v1/audio/transcriptions Transcription model
Create an illustration POST /v1/images/generations Image-generation model
Read text aloud POST /v1/audio/speech Supported speech-output runtime and voice

These are separate capabilities. A text-only model does not become a vision model because the request contains an image.

Use a Mac for the complete set in this guide. The verified Windows stable package does not include the same local speech-output runtime. Windows can use supported local text/vision, transcription and image-generation routes, but check the installed model and runtime for each task.

Core inference and the gateway do not require Pro capture. Start with one capability before building a UI around several at once.

In OGAD, download and select a supported local vision model in Models. Open Gateway and check the local address. The example below uses the normal 7878 port; replace it if your app shows another one.

Save a small PNG image as example.png. This Python 3 script sends the local image as a data URL, so it does not need a remotely hosted image:

import base64
import json
from pathlib import Path
from urllib.request import Request, urlopen

BASE = "http://127.0.0.1:7878"

def request_json(path, payload=None):
    data = None if payload is None else json.dumps(payload).encode("utf-8")
    request = Request(BASE + path, data=data,
                      headers={"Content-Type": "application/json"})
    with urlopen(request, timeout=240) as response:
        return json.load(response)

models = request_json("/v1/models")["data"]
model = next((item for item in models
              if item.get("kind") == "vision" and not item.get("remote")), None)
if model is None:
    raise SystemExit("Select a downloaded local vision model in OGAD first.")

encoded = base64.b64encode(Path("example.png").read_bytes()).decode("ascii")
result = request_json("/v1/chat/completions", {
    "model": model["id"],
    "messages": [{"role": "user", "content": [
        {"type": "text", "text": "Describe the visible image. Mark unclear details."},
        {"type": "image_url", "image_url": {"url": "data:image/png;base64," + encoded}},
    ]}],
    "max_tokens": 160,
    "stream": False,
})
print(result["choices"][0]["message"]["content"])

Compare the description with the source image. Image understanding can make mistakes, especially with small text or ambiguous details. Your app should let the user check the image beside the answer.

For image generation, select a downloaded image model and send a request such as:

{
  "prompt": "A simple watercolor illustration of a quiet reading desk, no text",
  "size": "512x512",
  "response_format": "b64_json"
}

Send it to /v1/images/generations. A successful JSON result contains an image in data[0].b64_json; decode it before displaying or saving it. Use a size supported by the chosen model. Do not assume a cloud image API's entire parameter set is supported here.

For spoken output on a supported Mac, prepare a local voice in OGAD first. The speech endpoint accepts JSON with input text and returns WAV audio by default. Your app must read those bytes as audio rather than trying to parse them as JSON.

Transcription uses a multipart file field, not a JSON string containing the recording's path. The running /docs reference gives each route's format.

Show useful and error states. The first request can include a model load, and a model that does not fit the available memory cannot be fixed by a longer HTTP timeout alone. Keep the first interaction small and avoid firing every model route at once.

Use local selections for every capability in an offline workflow. One configured remote provider can make an otherwise local app depend on internet. Complete first-use asset downloads before the offline check.

The gateway listens on network interfaces and its inference endpoints do not require an API key. These examples use 127.0.0.1 on the same computer. Keep the host on a trusted network and do not expose this port to the public internet.

These API routes are present in OGAD 0.0.51. The running gateway also serves its API reference at /docs.

Download OGAD and connect your app to one image question, one recording or one generated illustration. Add the next capability after the first gives a result the user can inspect.

── more in #ai-tools 4 stories · sorted by recency
── more on @ogad 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/how-to-add-offline-s…] indexed:0 read:4min 2026-09-29 · —