How I Built a Unified API Gateway for 200+ AI Models (Architecture Deep Dive) A developer built SarangAI, a unified AI gateway that routes requests to more than 200 models through a single OpenAI-compatible endpoint. The system uses a stateless router, per-provider adapters that normalize responses and streaming to OpenAI's SSE format, and a prepaid IDR billing layer backed by Redis for fast balance checks and a database for audit trails, plus a companion sarangai-cli terminal tool. Every developer who has worked with more than one AI provider knows the pain: If you want to use 4 different models, you have to manage 4 accounts, 4 invoices, and 4 sets of documentation. This isn't just annoying — it's a bottleneck that stops developers from experimenting with new models. I built SarangAI to solve this. It's a unified AI gateway that routes all requests to 200+ models through a single OpenAI-compatible endpoint . In this article, I'll walk through the architecture behind it, the design decisions I made, and the technical challenges that came up. At a high level, SarangAI consists of 4 main components: The flow is simple: Client → API Gateway → Router → Provider Adapter → AI Provider ↓ Billing & Rate Limiter This is the most important design decision I made. When I started building SarangAI, I had two options: Option 1: Build my own API format, my own docs, my own SDK. Option 2: Use the OpenAI format, which has become the de-facto standard. I chose option 2 . Here's why: This is what makes SarangAI usable in minutes, not hours. Every provider has a different response format. OpenAI has choices 0 .message.content , Anthropic has content 0 .text , Google has yet another structure. The solution: adapter pattern . Each provider has an adapter that: This keeps client-side code clean — they don't need to know which provider is being used. Streaming responses are tricky. Every provider sends chunks differently: data: {...} content block delta , etc. In SarangAI, I normalize all streaming to the same SSE format as OpenAI. So clients only need to handle one streaming format. Since SarangAI uses a prepaid IDR top-up model, I need to: For this, I use a combination of Redis for fast balance checks and a database for audit trails . One of SarangAI's main features is instant model switching . Users can change models without restarting their app. This means the router has to: I made this router stateless , so it can scale horizontally without issues. sarangai-cli Besides the API gateway, I also built a CLI tool that works directly from the terminal: npm install -g sarangai-cli sarang The CLI connects to the SarangAI endpoint and gives you an interactive workspace. You can: This is especially useful for developers who live in the terminal. 1. Standards matter. Choosing the OpenAI-compatible format was the best decision I made. It's what makes adoption fast. 2. Adapter patterns save lives. Without clean adapters, adding a new provider would be a nightmare. 3. Prepaid Subscription for developer tools. Developers hate monthly subscriptions. Prepaid gives them a sense of full control. 4. Documentation is a feature. No matter how good your architecture is, if the docs are bad, nobody will use it. If you work with multiple AI models regularly, give SarangAI a try: Install the CLI: npm install -g sarangai-cli I'm curious: what's your current setup for handling multiple AI providers? Do you use a library? Or manage them one by one? Share in the comments - I'd love to hear how other developers handle this.