Why I put one thin layer between my app and every LLM provider A developer described building a thin abstraction layer between their application and multiple LLM providers after finding API keys scattered across three places, duplicated retry logic, and no visibility into per-feature token costs. The layer routes requests by task type rather than by provider, though the developer notes tool-calling and structured-output differences still leak through and require per-provider adapter code. They credit per-task logging with revealing that a single feature consumed most of the tokens, and advise skipping the abstraction until a second provider is genuinely needed. On my first AI feature I called one provider straight from the code that needed it. It worked, and for a while that was the right call. Then I wanted a cheaper model for the boring tasks, like tagging and short summaries. That's when I noticed the provider was everywhere: keys in three places, retry logic copied around, and no clear idea what each feature actually cost. Chalk's post this week cited an a16z survey where 81% of CIOs at Global 2000 companies now use three or more model families. I'm nowhere near that scale, but even with two providers the mess shows up fast. It's small. Roughly: summarize or classify , not per provider. Another thing to maintain. Tool calling and structured output formats differ between providers, so the abstraction leaks. I ended up with small adapter code per provider anyway. It also tempts you to over build. I'd skip it until a second provider is actually on the table. The per task log. Seeing that one feature used most of the tokens changed what I worked on next more than any routing trick did. If you're running more than one model, did you build your own layer or use a gateway?