On my first AI feature I called one provider straight from the code that needed it. It worked, and for a while that was the right call.
Then I wanted a cheaper model for the boring tasks, like tagging and short summaries. That's when I noticed the provider was everywhere: keys in three places, retry logic copied around, and no clear idea what each feature actually cost.
Chalk's post this week cited an a16z survey where 81% of CIOs at Global 2000 companies now use three or more model families. I'm nowhere near that scale, but even with two providers the mess shows up fast.
It's small. Roughly:
summarize or classify, not per provider.
Another thing to maintain. Tool calling and structured output formats differ between providers, so the abstraction leaks. I ended up with small adapter code per provider anyway.
It also tempts you to over build. I'd skip it until a second provider is actually on the table.
The per task log. Seeing that one feature used most of the tokens changed what I worked on next more than any routing trick did.
If you're running more than one model, did you build your own layer or use a gateway?