{"slug": "why-i-put-one-thin-layer-between-my-app-and-every-llm-provider", "title": "Why I put one thin layer between my app and every LLM provider", "summary": "A developer described building a thin abstraction layer between their application and multiple LLM providers after finding API keys scattered across three places, duplicated retry logic, and no visibility into per-feature token costs. The layer routes requests by task type rather than by provider, though the developer notes tool-calling and structured-output differences still leak through and require per-provider adapter code. They credit per-task logging with revealing that a single feature consumed most of the tokens, and advise skipping the abstraction until a second provider is genuinely needed.", "body_md": "On my first AI feature I called one provider straight from the code that needed it. It worked, and for a while that was the right call.\n\nThen I wanted a cheaper model for the boring tasks, like tagging and short summaries. That's when I noticed the provider was everywhere: keys in three places, retry logic copied around, and no clear idea what each feature actually cost.\n\nChalk's post this week cited an a16z survey where 81% of CIOs at Global 2000 companies now use three or more model families. I'm nowhere near that scale, but even with two providers the mess shows up fast.\n\nIt's small. Roughly:\n\n`summarize` or `classify`, not per provider.\nAnother thing to maintain. Tool calling and structured output formats differ between providers, so the abstraction leaks. I ended up with small adapter code per provider anyway.\n\nIt also tempts you to over build. I'd skip it until a second provider is actually on the table.\n\nThe per task log. Seeing that one feature used most of the tokens changed what I worked on next more than any routing trick did.\n\nIf you're running more than one model, did you build your own layer or use a gateway?", "url": "https://wpnews.pro/news/why-i-put-one-thin-layer-between-my-app-and-every-llm-provider", "canonical_source": "https://dev.to/rishita_sharma_b0aa1ff81a/why-i-put-one-thin-layer-between-my-app-and-every-llm-provider-3a61", "published_at": "2026-10-07 04:37:48+00:00", "updated_at": "2026-10-07 04:47:31.483334+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "ai-tools", "mlops", "developer-tools"], "entities": ["a16z"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/why-i-put-one-thin-layer-between-my-app-and-every-llm-provider", "markdown": "https://wpnews.pro/news/why-i-put-one-thin-layer-between-my-app-and-every-llm-provider.md", "text": "https://wpnews.pro/news/why-i-put-one-thin-layer-between-my-app-and-every-llm-provider.txt", "jsonld": "https://wpnews.pro/news/why-i-put-one-thin-layer-between-my-app-and-every-llm-provider.jsonld"}}