{"slug": "openai-compatible-apis-vs-provider-abstraction-what-actually-transfers-across", "title": "OpenAI-compatible APIs vs provider abstraction: What actually transfers across LLM providers?", "summary": "A developer built llm-api-adapter, an open-source Python SDK that provides a provider-neutral contract across OpenAI, Anthropic, Google, Mistral, xAI, Qwen, Kimi, DeepSeek, and Z.ai, to address portability gaps between OpenAI-compatible APIs. The writeup documents that compatibility operates on a ladder from transport to operational levels, and that Anthropic's OpenAI compatibility endpoint hoists and concatenates system and developer messages into a single initial system message, so instruction placement and role distinctions are not preserved even though the request succeeds.", "body_md": "OpenAI-compatible APIs promise something very attractive: write one integration, change the `base_url` and model, and move between LLM providers without rewriting your application.\n\nFor basic chat, this often works surprisingly well. **The more interesting question is what that compatibility actually guarantees.**\n\nThe HTTP format? The accepted parameters? Or the behavior your application can rely on?\n\nI ran into this distinction repeatedly while building [**llm-api-adapter**](https://github.com/Inozem/llm_api_adapter), an open-source, lightweight Python SDK that provides a provider-neutral contract across OpenAI, Anthropic, Google, Mistral, xAI, Qwen, Kimi, DeepSeek, and Z.ai.\n\nI’ll use it throughout this article as one example of how these portability problems can be handled. The point is not that every application needs another abstraction layer. In many cases, an OpenAI-compatible endpoint is enough.\n\nThe goal is to find where it stops being enough.\n\nFor this comparison, I’m staying mostly inside the **chat contract**: instruction semantics, reasoning controls, and unsupported or model-specific parameters. Tool calling and structured outputs introduce another set of differences and deserve separate comparisons.\n\n**Verification note:** Provider documentation was checked on October 8, 2026. Library examples use `llm-api-adapter` 0.9.8 and follow its tested public API.\n\nEach provider example assumes an OpenAI client configured with that provider’s API key and base URL. I use names such as `anthropic_client`, `gemini_client`, and `qwen_client` below to make that explicit.\n\nAt the simplest level, compatibility looks like this:\n\n``` python\nfrom openai import OpenAI\n\nclient = OpenAI(\n    api_key=PROVIDER_API_KEY,\n    base_url=PROVIDER_BASE_URL,\n)\n\nresponse = client.chat.completions.create(\n    model=MODEL,\n    messages=[\n        {\n            \"role\": \"user\",\n            \"content\": \"Why do database indexes speed up reads?\"\n        }\n    ],\n)\n\nprint(response.choices[0].message.content)\n```\n\nFor a simple user-message → text-response workflow, that may be all you need.\n\nBut production applications tend to depend on more: instruction roles, reasoning controls, usage metadata, and failure behavior.\n\nI find it useful to think about compatibility as a ladder:\n\n```\ntransport compatibility\n        ↓\nrequest-shape compatibility\n        ↓\nfeature compatibility\n        ↓\nsemantic compatibility\n        ↓\noperational compatibility\n```\n\nTwo APIs can be compatible near the top and still require different application logic further down.\n\nConsider an OpenAI-style conversation:\n\n```\nmessages = [\n    {\n        \"role\": \"system\",\n        \"content\": \"Follow company policy A.\"\n    },\n    {\n        \"role\": \"user\",\n        \"content\": \"First question\"\n    },\n    {\n        \"role\": \"developer\",\n        \"content\": \"For the next step, prefer policy B.\"\n    },\n    {\n        \"role\": \"user\",\n        \"content\": \"Second question\"\n    },\n]\n```\n\nAnthropic's OpenAI compatibility endpoint accepts that shape:\n\n```\nresponse = anthropic_client.chat.completions.create(\n    model=\"claude-opus-5-5\",\n    messages=messages,\n)\n```\n\nBut the compatibility layer documents a transformation: `system` and `developer` messages are hoisted to the beginning and concatenated into one initial system message. It also documents a number of OpenAI parameters that are accepted but ignored. [Anthropic: OpenAI SDK compatibility](https://platform.claude.com/docs/en/cli-sdks-libraries/libraries/openai-sdk)\n\nConceptually, this:\n\n```\nsystem: Follow company policy A.\n\nuser: First question\n\ndeveloper: For the next step, prefer policy B.\n\nuser: Second question\n```\n\nis represented more like:\n\n```\nsystem:\nFollow company policy A.\nFor the next step, prefer policy B.\n\nuser: First question\n\nuser: Second question\n```\n\nThe request succeeds, but **instruction placement and role distinctions are not preserved**.\n\nThis is specifically a property of Anthropic's OpenAI compatibility layer. The native Claude API supports mid-conversation `system` messages on selected models, including Opus 5.5, subject to documented placement constraints. [Anthropic: Mid-conversation system messages](https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages)\n\nIf the requirement can be expressed as a portable initial instruction, the application can use `Prompt`:\n\n```\nfrom llm_api_adapter.models.messages.chat_message import (\n    Prompt,\n    UserMessage,\n)\n\nmessages = [\n    Prompt(\"Follow company policy A.\"),\n    UserMessage(\"First question\"),\n    UserMessage(\"Second question\"),\n]\n\nresponse = adapter.chat(messages=messages)\n```\n\nIf the application specifically depends on a mid-conversation OpenAI `developer` role, that distinction is not treated as universally portable. It should remain provider-specific instead of being silently translated into something that means something else.\n\nSometimes the correct normalization is to expose a boundary rather than hide it.\n\n`reasoning_effort` parameter can mean several different things\nReasoning controls are a useful example because several providers deliberately expose OpenAI-style parameters.\n\nSuppose the application asks for:\n\n```\nreasoning_effort=\"high\"\n```\n\nThat looks portable. Its meaning is not.\n\nAnthropic's compatibility endpoint accepts:\n\n```\nresponse = anthropic_client.chat.completions.create(\n    model=\"claude-opus-5-5\",\n    messages=messages,\n    reasoning_effort=\"high\",\n)\n```\n\nBut Anthropic documents `reasoning_effort` as ignored by this compatibility layer. Claude has native mechanisms for controlling reasoning behavior, and Anthropic recommends the native API when applications need access to the full thinking feature set. [Anthropic: OpenAI SDK compatibility](https://platform.claude.com/docs/en/cli-sdks-libraries/libraries/openai-sdk)\n\nGoogle takes a different approach:\n\n```\nresponse = gemini_client.chat.completions.create(\n    model=\"gemini-3.6-flash\",\n    messages=messages,\n    reasoning_effort=\"high\",\n)\n```\n\nThis example uses the stable `gemini-3.6-flash` model. Google's OpenAI compatibility layer maps `reasoning_effort` onto Gemini's own thinking controls rather than simply ignoring it. [Google: OpenAI compatibility](https://ai.google.dev/gemini-api/docs/openai)\n\nQwen 3.8 adds another interpretation:\n\n```\nresponse = qwen_client.chat.completions.create(\n    model=\"qwen3.8-flash\",\n    messages=messages,\n    reasoning_effort=\"high\",\n)\n```\n\nQwen 3.8 uses `low`, `medium`, and `xhigh` as its native reasoning levels, with `xhigh` as the default. OpenAI-compatible values are mapped onto that scale: for example, `minimal` maps to `low`, while `high` and `max` map to `xhigh`. `none` disables thinking.\n\nInvalid values can cause an error, and `reasoning_effort` cannot be supplied together with `thinking_budget`. [Qwen: OpenAI-compatible Chat API](https://docs.qwencloud.com/api-reference/chat/openai-chat)\n\nSo the same-looking setting can mean:\n\n```\nClaude\nhigh → ignored by compatibility layer\n\nGemini\nhigh → translated to Gemini thinking controls\n\nQwen\nhigh → mapped to xhigh\n```\n\nThe application uses one setting:\n\n```\nresponse = adapter.chat(\n    messages=[\n        UserMessage(\"Solve this architecture problem.\")\n    ],\n    reasoning_level=\"high\",\n)\n```\n\n`reasoning_level` expresses configuration intent. For categorical models, the adapter projects portable levels onto the model's registered native levels. For models with numeric reasoning budgets, it computes corresponding values within the registered supported range.\n\nPortable levels are:\n\n```\nnone\nminimal\nlow\nmedium\nhigh\nvery_high\n```\n\nProvider-native values such as `xhigh` can still be used for models that support them, but they are intentionally not considered portable.\n\n**These mappings normalize configuration intent; they do not guarantee equivalent reasoning depth, quality, latency, or cost across models.**\n\nThe abstraction can normalize how the application asks for reasoning. It cannot make different models reason equivalently.\n\nCompatibility layers often ignore unsupported fields instead of failing the request. That is useful for migrations, but it changes what a successful response proves.\n\nFor Anthropic:\n\n```\nresponse = anthropic_client.chat.completions.create(\n    model=\"claude-opus-5-5\",\n    messages=messages,\n    reasoning_effort=\"high\",\n    service_tier=\"priority\",\n)\n```\n\nAnthropic currently documents both fields as ignored through its OpenAI compatibility layer. [Anthropic: OpenAI SDK compatibility](https://platform.claude.com/docs/en/cli-sdks-libraries/libraries/openai-sdk)\n\nDeepSeek follows a similar approach:\n\n```\nresponse = deepseek_client.responses.create(\n    model=\"deepseek-flash\",\n    input=\"Explain MVCC.\",\n    metadata={\n        \"request_id\": \"abc-123\"\n    },\n    service_tier=\"priority\",\n)\n```\n\n`metadata` and `service_tier` are unsupported, and DeepSeek states that unsupported Responses parameters are silently ignored so existing clients can connect without modification. [DeepSeek: Using the Responses API](https://api-docs.deepseek.com/guides/responses_api/)\n\nSo this:\n\n```\nassert response is not None\n```\n\nproves that the provider accepted the request. It does not prove that every requested behavior was applied.\n\n`llm-api-adapter` keeps verified model-specific compatibility rules in its registry.\n\nFor example, OpenAI documents that when GPT-6 reasoning effort is above `none`, sampling parameters including `temperature` and `top_p` must be removed. [OpenAI: Using GPT-6](https://developers.openai.com/api/docs/guides/latest-model)\n\nApplication code can still contain:\n\n```\nresponse = adapter.chat(\n    messages=[\n        UserMessage(\n            \"Compare optimistic and pessimistic locking.\"\n        )\n    ],\n    reasoning_level=\"high\",\n    temperature=0.2,\n)\n```\n\nFor `gpt-6-sol`, the registered request rule removes `temperature` before the request is sent.\n\nWhen a compatibility rule removes a parameter, the adapter warns if its value differs from the registered default. In this example, `temperature=0.2` triggers a warning.\n\nThe implementation lives in the adapter's model request-rule layer: [request_rules.py](https://github.com/Inozem/llm_api_adapter/blob/main/src/llm_api_adapter/llm_registry/request_rules.py).\n\nThe goal is not to manufacture feature parity. It is to make known incompatibilities predictable instead of forcing application code to infer them from successful HTTP requests.\n\nYou can solve every example above directly:\n\n```\nif provider == \"anthropic\":\n    ...\n\nelif provider == \"google\":\n    ...\n\nelif provider == \"qwen\":\n    ...\n```\n\nBut once application code knows which provider ignores a field, which one maps it, which one uses a different reasoning scale, and which model requires parameters to be transformed or omitted, it already contains a provider abstraction — just distributed through conditionals.\n\nA dedicated layer moves that knowledge behind one boundary:\n\n```\napplication\n    ↓\nportable application contract\n    ↓\nmodel capability rules\n    ↓\nprovider-specific translation\n    ↓\nprovider API\n```\n\nFor `llm-api-adapter`, the shared layer includes typed messages and responses, reasoning intent, streaming behavior, usage/cost information, and normalized error categories. Provider implementations own wire formats, request rules, reasoning mappings, response parsing, and stream-event differences.\n\nThe goal is not to make providers look identical. It is to keep differences out of application code where a stable mapping actually exists.\n\nAfter that provider-specific work has been moved behind the boundary, the application call is much smaller:\n\n```\nfrom llm_api_adapter.models.messages.chat_message import (\n    Prompt,\n    UserMessage,\n)\nfrom llm_api_adapter.universal_adapter import (\n    UniversalLLMAPIAdapter,\n)\n\nmessages = [\n    Prompt(\"Answer as a backend engineer.\"),\n    UserMessage(\n        \"Explain connection pooling in two paragraphs.\"\n    ),\n]\n\nadapter = UniversalLLMAPIAdapter(\n    organization=organization,\n    model=model,\n    api_key=api_key,\n)\n\nresponse = adapter.chat(\n    messages=messages,\n    max_tokens=400,\n    reasoning_level=\"medium\",\n    **provider_call_config,\n)\n\nprint(response.content)\n```\n\nFor most integrations:\n\n```\nprovider_call_config = {}\n```\n\nQwen is one example where a provider-specific call parameter remains necessary:\n\n```\nprovider_call_config = {\n    \"workspace_id\": QWEN_WORKSPACE_ID,\n}\n```\n\nThe current Qwen package uses Alibaba Model Studio's Frankfurt Global deployment through its Anthropic-compatible Messages endpoint, which requires `workspace_id` on each call. [Qwen integration README](https://github.com/Inozem/llm_api_adapter/blob/main/packages/organizations/qwen/README.md)\n\nThat requirement stays explicit, while the rest of the application call remains unchanged.\n\nI still would not add an abstraction layer to every application.\n\nAn OpenAI-compatible endpoint can be the simpler choice when you already have an OpenAI integration, want to evaluate another model quickly, mostly need basic chat, and have tested the subset of the API your application depends on.\n\nIf your product is built deeply around one provider's unique capabilities, its native API can be the better abstraction.\n\nA provider layer becomes more useful when **multi-provider support itself is an application requirement**. Then the important question is no longer whether another provider accepts the same JSON. It is whether the same application assumptions still hold.\n\n**OpenAI-compatible APIs give you protocol portability. A provider abstraction can give you application-contract portability.**\n\nThose layers overlap, but they are not the same thing.\n\nThe examples above use `llm-api-adapter` **0.9.8**. The project is open source and currently supports OpenAI, Anthropic, Google, Mistral, xAI, Qwen, Kimi, DeepSeek, and Z.ai through a mix of native and compatibility APIs.", "url": "https://wpnews.pro/news/openai-compatible-apis-vs-provider-abstraction-what-actually-transfers-across", "canonical_source": "https://dev.to/inozem/openai-compatible-apis-vs-provider-abstraction-what-actually-transfers-across-llm-providers-2afd", "published_at": "2026-10-08 17:08:24+00:00", "updated_at": "2026-10-08 17:20:51.350743+00:00", "lang": "en", "topics": ["ai-tools", "developer-tools", "large-language-models", "ai-agents"], "entities": ["OpenAI", "Anthropic", "Google", "Mistral", "xAI", "Qwen", "Kimi", "DeepSeek"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/openai-compatible-apis-vs-provider-abstraction-what-actually-transfers-across", "markdown": "https://wpnews.pro/news/openai-compatible-apis-vs-provider-abstraction-what-actually-transfers-across.md", "text": "https://wpnews.pro/news/openai-compatible-apis-vs-provider-abstraction-what-actually-transfers-across.txt", "jsonld": "https://wpnews.pro/news/openai-compatible-apis-vs-provider-abstraction-what-actually-transfers-across.jsonld"}}