If you're building with LLMs, the model itself is often the easiest part.
The annoying part comes later.
You start with one provider. Then you want to test another model. Soon you have multiple API keys, different pricing structures, different endpoints, separate billing accounts, and provider-specific code scattered across your project.
I ran into exactly this problem while building AI applications.
Suppose an application needs access to several model families:
Using each provider directly can mean maintaining several integrations.
Even when APIs look similar, authentication, model names, endpoints, billing and availability can differ.
There is also a second problem: cost.
For experiments, agents and applications processing large numbers of tokens, API costs can become significant surprisingly quickly.
So I wanted two things:
A useful approach is to standardize around the OpenAI API format.
Instead of changing application logic whenever you change providers, you keep essentially the same request structure and change the base URL and model.
For example, an application using the OpenAI Python SDK can look roughly like this:
from openai import OpenAI
client = OpenAI(
api_key="YOUR_API_KEY",
base_url="YOUR_OPENAI_COMPATIBLE_ENDPOINT"
)
response = client.chat.completions.create(
model="YOUR_MODEL",
messages=[
{
"role": "user",
"content": "Explain why API compatibility matters."
}
]
)
print(response.choices[0].message.content)
The important part isn't the few lines of code.
It's that the application no longer needs to be tightly coupled to a single model provider.
While working on this problem, I built MandAPI, an OpenAI-compatible multi-model API gateway.
The idea is deliberately simple:
one API format, one API key, multiple AI model families.
Instead of maintaining completely separate integrations, developers can access models from families such as GPT, Claude, Gemini and DeepSeek through an OpenAI-compatible interface.
That also makes price experimentation much easier.
If a workload doesn't require the most expensive model, you can move it to a cheaper model without redesigning the application.
A common mistake is using the most capable model for every request.
In a real application, workloads are usually mixed.
Some requests need strong reasoning.
Others are simple extraction, classification, rewriting, summarization or conversational tasks.
A better architecture can look like:
Complex reasoning
↓
High-capability model
Normal generation
↓
Mid-cost model
Simple/high-volume tasks
↓
Low-cost model
Once multiple models share a compatible interface, routing workloads this way becomes much easier.
For high-token applications, the difference can become substantial.
Token price is important, but it isn't the only variable.
When comparing AI APIs, I now look at:
The cheapest model on paper isn't necessarily the cheapest model for the application.
A model that needs twice as many attempts to produce an acceptable answer may actually cost more.
There is another benefit that is easy to underestimate: avoiding provider lock-in.
Your application talks to an interface rather than being designed around one specific provider.
That makes it easier to:
This becomes increasingly useful because AI model pricing and capabilities change very quickly.
I'm continuing to work on MandAPI as both an API gateway and a model/pricing research project.
One area I'm particularly interested in is transparent LLM price comparison.
I've started publishing some of the underlying developer resources and pricing research openly on GitHub:
The goal is to make it easier for developers to compare models based on actual API economics rather than marketing pages.
I'm also especially interested in making AI APIs easier to purchase and use in markets where international billing can be inconvenient, including Brazil.
You don't necessarily need to rewrite your AI application to experiment with different providers.
An OpenAI-compatible abstraction layer can make model switching surprisingly simple.
And once switching becomes simple, price becomes something you can optimize continuously instead of something you're locked into.
If you're building an AI product with significant API usage, it's worth designing for model portability from the beginning.
I'd be interested to hear how other developers are handling multi-model routing and API costs.