This is the engineering summary of a full reference architecture. The paper, its object model as JSON and the model register are open: https://muhammadumar89.github.io/codeninja-research/sovereign-hse-pakistan/.
An oil and gas operator in Pakistan wants one thing from AI in its health, safety and environment department: a warning before the next incident, not a report after it. The data to do that already exists, spread across SAP EHS, SCADA and fire-and-gas historians, camera feeds and scanned investigation files. The constraint is just as clear. None of it may leave the operator's own infrastructure, and no third-party AI API may sit in the serving path.
Here is how that constraint turns into hardware, models and money.
On an air-gapped platform you cannot call a hosted model, so every model must be self-hosted, and the licence decides whether the operator owns what it runs. Every pick lets the operator hold, run and fine-tune the weights inside its own boundary:
| Role | Model | Licence |
|---|---|---|
| Reasoning, cited answers, agents | GLM 5.3 open weights, 753B mixture-of-experts at FP8 | bespoke; purely internal use is exempt from its managed-service review |
| Time-series anomaly and early warning | amazon/chronos-2 | Apache-2.0 |
| Vision detection baseline | Roboflow/rf-detr-large (Nano to Large only) | Apache-2.0 |
| Tracking across frames | Roboflow trackers | Apache-2.0 |
| Multilingual retrieval (English, Urdu, Roman Urdu) | BAAI/bge-m3 | MIT |
| OCR of scanned permits and reports | PaddlePaddle/PaddleOCR-VL-1.6 | Apache-2.0 |
RF-DETR's larger checkpoints ship under a different platform licence, so the design stops at Large.
Take the largest filed parameter count, multiply by bytes per parameter at the serving precision, then add a planning factor for the KV cache and activations so long incident histories fit:
weights = parameters x bytes per parameter
need = weights x 1.2 planning factor
nodes = ceil(need / (cards per node x memory per card))
GLM 5.3 is filed at 753B parameters. At FP8 that is 753 GB of weights and 904 GB with headroom, so one node of eight 141 GB cards (1,128 GB) holds it, leaving 375 GB beside the weights for KV cache: long incident histories and concurrent users.
Detection and forecasting must keep up with cameras and sensors even if the link to the central tier drops. Detection runs on edge nodes inside the plant network, reusing the operator's NPU compute where it exists. Forecasting, OCR and embeddings run on a site inference server beside the historian: Chronos-2, PaddleOCR-VL 1.6 and BGE-M3 together weigh under 4 GB. Edge and site compute are sized by stream and decode load, not by model count.
Three years, public prices, electricity at Pakistan's B3 industrial tariff:
| Option | Three-year cost (USD) |
|---|---|
| Own the hardware, with support and power | about 670,000 |
| Rent the same GPUs, AWS UAE region, three-year plan | 1.85 million |
| Rent the same GPUs, AWS UAE region, on demand | 2.82 million |
| Buy a closed frontier model by the token, 50 users | 1.04 to 2.95 million |
No hyperscaler runs a region inside Pakistan, so every rented option also moves the data abroad. The full workings and sources are in Appendix A of the paper.
141 GB HBM-class accelerators need a US export licence for Pakistan (Country Group D:4). Approved channels have delivered thousands of GPUs to Pakistani operators, and the design's first phase confirms the installed inventory before anything is bought, with a fallback to a mid-size model on existing hardware.
The object model ships as JSON in the repository in a format meant for import into an ontology platform, with every object's anchor system, properties, status vocabulary and links. Take it, change it, cite it.
Umar Bilal, Cofounder of CodeNinja. CodeNinja is a Middle Eastern American artificial intelligence lab that puts autonomy in physical operations.