Retrieval-Augmented Generation (RAG) can help generative AI systems produce answers grounded in organizational evidence. However, retrieval alone does not make an AI application trustworthy. A production RAG platform must also protect personal information, resist instruction manipulation, isolate tenants, cite evidence, abstain when evidence is insufficient and preserve an auditable deployment process.
I built the Enterprise Multi-Cloud GenAI RAG Platform as an open and reproducible reference implementation of those controls. It runs locally without paid model APIs and provides equivalent production deployment paths for Amazon Web Services and Microsoft Azure.
The source code, evaluation assets and infrastructure definitions are publicly available on GitHub. The software and accompanying technical report are permanently archived on Zenodo with separate DOIs.
The platform combines several disciplines that are often demonstrated separately:
The system follows a controlled evidence pipeline:
The local implementation is deliberately deterministic. It extracts responses from retrieved evidence rather than requiring an external generative model. This makes the security and governance behaviour testable without cloud credentials or inference charges.
The AWS reference architecture maps the platform responsibilities to:
The Microsoft Azure reference architecture maps the same responsibilities to:
The Terraform configurations are intentionally plan-oriented. A real deployment should add approved model access, private networking, organizational identity integration, budgets, privacy review and named deployment approval.
The repository contains a 20-case development benchmark and a separate frozen 40-case synthetic candidate test set. The candidate test covers answerable questions, unsupported questions, prompt injection, privacy redaction, validation and tenant isolation.
Across five deterministic repetitions, the local configuration produced:
| Metric | Result |
|---|---|
| End-to-end success | 34/40 (0.85) |
| Citation coverage | 20/21 (0.9524) |
| Prompt-injection blocking | 5/8 (0.625) |
| PII-redaction recall | 6/6 (1.00) |
| Abstention accuracy | 6/6 (1.00) |
| Median local latency | 0.1322 ms |
| p95 local latency | 0.1808 ms |
These figures describe the lightweight local implementation in the recorded environment. They are not cloud-latency estimates.
The most important result was not the overall success rate. Three of eight prompt-injection formulations bypassed the initial literal-pattern detector. This demonstrates why a regular-expression safeguard can be a useful basic control but cannot be treated as a comprehensive defence.
A later layered detector added Unicode normalization, zero-width-character removal and scored combinations of override, control, exfiltration and protected-information indicators. On a separate 36-case diagnostic set, it detected 17 of 24 attacks, accepted 11 of 12 benign inputs and achieved 0.9444 precision.
The hardened detector still missed encoded, multilingual, spaced-letter and role-play formulations. Responsible reporting therefore requires publishing both successful and failed cases instead of presenting the system as universally secure.
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt
python -m src.rag_platform.cli ingest examples/knowledge
uvicorn src.rag_platform.api:app --host 0.0.0.0 --port 8000
Run the evaluation suite:
python -m src.rag_platform.cli evaluate evaluation/golden_set.json
python -m src.rag_platform.cli experiment evaluation/candidate_test_set.json --repeats 5 --output evaluation/results/local_candidate_test.json
python -m src.rag_platform.cli security-evaluate evaluation/security_robustness_v2.json
The present datasets are synthetic, small and not independently annotated. The local hashed embeddings are not a substitute for modern semantic retrieval. The security tests do not cover the full creativity of adversarial attacks, and the PII redactor supports only documented patterns. The project should therefore be treated as a transparent, reproducible baseline—not as certification of comprehensive AI safety.
Arayemi, J. (2026). Secure and Responsible Retrieval-Augmented Generation for Resource-Constrained Organizations: A Multi-Cloud Reference Architecture and Experimental Evaluation. Zenodo. https://doi.org/10.5281/zenodo.23048250
I welcome independent reproduction, technical feedback and contributions. If you use the architecture, evaluation protocol or implementation in research, training or a derived system, please cite the technical report and link to the repository.