cd /news/ai-infrastructure/sizing-a-sovereign-air-gapped-ai-sta… · home › topics › ai-infrastructure › article
[ARTICLE · art-144331] src=dev.to ↗ pub= topic=ai-infrastructure verified=true sentiment=· neutral

Sizing a sovereign, air-gapped AI stack for oil and gas health, safety and environment (HSE) in Pakistan

An engineer published a reference architecture for a sovereign, air-gapped AI stack for oil and gas health, safety and environment (HSE) work in Pakistan, sizing self-hosted models including GLM 5.3, Amazon Chronos-2, RF-DETR, BGE-M3 and PaddleOCR-VL 1.6 onto operator-owned hardware. The design puts a 753B-parameter FP8 GLM 5.3 on a single node of eight 141 GB accelerators and estimates a three-year cost of about $670,000 for owned hardware versus $1.85 million to $2.95 million for rented or token-priced alternatives, noting that 141 GB HBM-class accelerators require a US export licence for Pakistan.

by read3 min views1 publishedOct 3, 2026

This is the engineering summary of a full reference architecture. The paper, its object model as JSON and the model register are open: https://muhammadumar89.github.io/codeninja-research/sovereign-hse-pakistan/.

An oil and gas operator in Pakistan wants one thing from AI in its health, safety and environment department: a warning before the next incident, not a report after it. The data to do that already exists, spread across SAP EHS, SCADA and fire-and-gas historians, camera feeds and scanned investigation files. The constraint is just as clear. None of it may leave the operator's own infrastructure, and no third-party AI API may sit in the serving path.

Here is how that constraint turns into hardware, models and money.

On an air-gapped platform you cannot call a hosted model, so every model must be self-hosted, and the licence decides whether the operator owns what it runs. Every pick lets the operator hold, run and fine-tune the weights inside its own boundary:

Role Model Licence
Reasoning, cited answers, agents GLM 5.3 open weights, 753B mixture-of-experts at FP8 bespoke; purely internal use is exempt from its managed-service review
Time-series anomaly and early warning amazon/chronos-2 Apache-2.0
Vision detection baseline Roboflow/rf-detr-large (Nano to Large only) Apache-2.0
Tracking across frames Roboflow trackers Apache-2.0
Multilingual retrieval (English, Urdu, Roman Urdu) BAAI/bge-m3 MIT
OCR of scanned permits and reports PaddlePaddle/PaddleOCR-VL-1.6 Apache-2.0

RF-DETR's larger checkpoints ship under a different platform licence, so the design stops at Large.

Take the largest filed parameter count, multiply by bytes per parameter at the serving precision, then add a planning factor for the KV cache and activations so long incident histories fit:

weights   = parameters x bytes per parameter
need      = weights x 1.2 planning factor
nodes     = ceil(need / (cards per node x memory per card))

GLM 5.3 is filed at 753B parameters. At FP8 that is 753 GB of weights and 904 GB with headroom, so one node of eight 141 GB cards (1,128 GB) holds it, leaving 375 GB beside the weights for KV cache: long incident histories and concurrent users.

Detection and forecasting must keep up with cameras and sensors even if the link to the central tier drops. Detection runs on edge nodes inside the plant network, reusing the operator's NPU compute where it exists. Forecasting, OCR and embeddings run on a site inference server beside the historian: Chronos-2, PaddleOCR-VL 1.6 and BGE-M3 together weigh under 4 GB. Edge and site compute are sized by stream and decode load, not by model count.

Three years, public prices, electricity at Pakistan's B3 industrial tariff:

Option Three-year cost (USD)
Own the hardware, with support and power about 670,000
Rent the same GPUs, AWS UAE region, three-year plan 1.85 million
Rent the same GPUs, AWS UAE region, on demand 2.82 million
Buy a closed frontier model by the token, 50 users 1.04 to 2.95 million

No hyperscaler runs a region inside Pakistan, so every rented option also moves the data abroad. The full workings and sources are in Appendix A of the paper.

141 GB HBM-class accelerators need a US export licence for Pakistan (Country Group D:4). Approved channels have delivered thousands of GPUs to Pakistani operators, and the design's first phase confirms the installed inventory before anything is bought, with a fallback to a mid-size model on existing hardware.

The object model ships as JSON in the repository in a format meant for import into an ontology platform, with every object's anchor system, properties, status vocabulary and links. Take it, change it, cite it.

Umar Bilal, Cofounder of CodeNinja. CodeNinja is a Middle Eastern American artificial intelligence lab that puts autonomy in physical operations.

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @glm 5.3 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/sizing-a-sovereign-a…] indexed:0 read:3min 2026-10-03 · —