# OpenAI-compatible APIs vs provider abstraction: What actually transfers across LLM providers?

> Source: <https://dev.to/inozem/openai-compatible-apis-vs-provider-abstraction-what-actually-transfers-across-llm-providers-2afd>
> Published: 2026-10-08 17:08:24+00:00

OpenAI-compatible APIs promise something very attractive: write one integration, change the `base_url` and model, and move between LLM providers without rewriting your application.

For basic chat, this often works surprisingly well. **The more interesting question is what that compatibility actually guarantees.**

The HTTP format? The accepted parameters? Or the behavior your application can rely on?

I ran into this distinction repeatedly while building [**llm-api-adapter**](https://github.com/Inozem/llm_api_adapter), an open-source, lightweight Python SDK that provides a provider-neutral contract across OpenAI, Anthropic, Google, Mistral, xAI, Qwen, Kimi, DeepSeek, and Z.ai.

I’ll use it throughout this article as one example of how these portability problems can be handled. The point is not that every application needs another abstraction layer. In many cases, an OpenAI-compatible endpoint is enough.

The goal is to find where it stops being enough.

For this comparison, I’m staying mostly inside the **chat contract**: instruction semantics, reasoning controls, and unsupported or model-specific parameters. Tool calling and structured outputs introduce another set of differences and deserve separate comparisons.

**Verification note:** Provider documentation was checked on October 8, 2026. Library examples use `llm-api-adapter` 0.9.8 and follow its tested public API.

Each provider example assumes an OpenAI client configured with that provider’s API key and base URL. I use names such as `anthropic_client`, `gemini_client`, and `qwen_client` below to make that explicit.

At the simplest level, compatibility looks like this:

``` python
from openai import OpenAI

client = OpenAI(
    api_key=PROVIDER_API_KEY,
    base_url=PROVIDER_BASE_URL,
)

response = client.chat.completions.create(
    model=MODEL,
    messages=[
        {
            "role": "user",
            "content": "Why do database indexes speed up reads?"
        }
    ],
)

print(response.choices[0].message.content)
```

For a simple user-message → text-response workflow, that may be all you need.

But production applications tend to depend on more: instruction roles, reasoning controls, usage metadata, and failure behavior.

I find it useful to think about compatibility as a ladder:

```
transport compatibility
        ↓
request-shape compatibility
        ↓
feature compatibility
        ↓
semantic compatibility
        ↓
operational compatibility
```

Two APIs can be compatible near the top and still require different application logic further down.

Consider an OpenAI-style conversation:

```
messages = [
    {
        "role": "system",
        "content": "Follow company policy A."
    },
    {
        "role": "user",
        "content": "First question"
    },
    {
        "role": "developer",
        "content": "For the next step, prefer policy B."
    },
    {
        "role": "user",
        "content": "Second question"
    },
]
```

Anthropic's OpenAI compatibility endpoint accepts that shape:

```
response = anthropic_client.chat.completions.create(
    model="claude-opus-5-5",
    messages=messages,
)
```

But the compatibility layer documents a transformation: `system` and `developer` messages are hoisted to the beginning and concatenated into one initial system message. It also documents a number of OpenAI parameters that are accepted but ignored. [Anthropic: OpenAI SDK compatibility](https://platform.claude.com/docs/en/cli-sdks-libraries/libraries/openai-sdk)

Conceptually, this:

```
system: Follow company policy A.

user: First question

developer: For the next step, prefer policy B.

user: Second question
```

is represented more like:

```
system:
Follow company policy A.
For the next step, prefer policy B.

user: First question

user: Second question
```

The request succeeds, but **instruction placement and role distinctions are not preserved**.

This is specifically a property of Anthropic's OpenAI compatibility layer. The native Claude API supports mid-conversation `system` messages on selected models, including Opus 5.5, subject to documented placement constraints. [Anthropic: Mid-conversation system messages](https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages)

If the requirement can be expressed as a portable initial instruction, the application can use `Prompt`:

```
from llm_api_adapter.models.messages.chat_message import (
    Prompt,
    UserMessage,
)

messages = [
    Prompt("Follow company policy A."),
    UserMessage("First question"),
    UserMessage("Second question"),
]

response = adapter.chat(messages=messages)
```

If the application specifically depends on a mid-conversation OpenAI `developer` role, that distinction is not treated as universally portable. It should remain provider-specific instead of being silently translated into something that means something else.

Sometimes the correct normalization is to expose a boundary rather than hide it.

`reasoning_effort` parameter can mean several different things
Reasoning controls are a useful example because several providers deliberately expose OpenAI-style parameters.

Suppose the application asks for:

```
reasoning_effort="high"
```

That looks portable. Its meaning is not.

Anthropic's compatibility endpoint accepts:

```
response = anthropic_client.chat.completions.create(
    model="claude-opus-5-5",
    messages=messages,
    reasoning_effort="high",
)
```

But Anthropic documents `reasoning_effort` as ignored by this compatibility layer. Claude has native mechanisms for controlling reasoning behavior, and Anthropic recommends the native API when applications need access to the full thinking feature set. [Anthropic: OpenAI SDK compatibility](https://platform.claude.com/docs/en/cli-sdks-libraries/libraries/openai-sdk)

Google takes a different approach:

```
response = gemini_client.chat.completions.create(
    model="gemini-3.6-flash",
    messages=messages,
    reasoning_effort="high",
)
```

This example uses the stable `gemini-3.6-flash` model. Google's OpenAI compatibility layer maps `reasoning_effort` onto Gemini's own thinking controls rather than simply ignoring it. [Google: OpenAI compatibility](https://ai.google.dev/gemini-api/docs/openai)

Qwen 3.8 adds another interpretation:

```
response = qwen_client.chat.completions.create(
    model="qwen3.8-flash",
    messages=messages,
    reasoning_effort="high",
)
```

Qwen 3.8 uses `low`, `medium`, and `xhigh` as its native reasoning levels, with `xhigh` as the default. OpenAI-compatible values are mapped onto that scale: for example, `minimal` maps to `low`, while `high` and `max` map to `xhigh`. `none` disables thinking.

Invalid values can cause an error, and `reasoning_effort` cannot be supplied together with `thinking_budget`. [Qwen: OpenAI-compatible Chat API](https://docs.qwencloud.com/api-reference/chat/openai-chat)

So the same-looking setting can mean:

```
Claude
high → ignored by compatibility layer

Gemini
high → translated to Gemini thinking controls

Qwen
high → mapped to xhigh
```

The application uses one setting:

```
response = adapter.chat(
    messages=[
        UserMessage("Solve this architecture problem.")
    ],
    reasoning_level="high",
)
```

`reasoning_level` expresses configuration intent. For categorical models, the adapter projects portable levels onto the model's registered native levels. For models with numeric reasoning budgets, it computes corresponding values within the registered supported range.

Portable levels are:

```
none
minimal
low
medium
high
very_high
```

Provider-native values such as `xhigh` can still be used for models that support them, but they are intentionally not considered portable.

**These mappings normalize configuration intent; they do not guarantee equivalent reasoning depth, quality, latency, or cost across models.**

The abstraction can normalize how the application asks for reasoning. It cannot make different models reason equivalently.

Compatibility layers often ignore unsupported fields instead of failing the request. That is useful for migrations, but it changes what a successful response proves.

For Anthropic:

```
response = anthropic_client.chat.completions.create(
    model="claude-opus-5-5",
    messages=messages,
    reasoning_effort="high",
    service_tier="priority",
)
```

Anthropic currently documents both fields as ignored through its OpenAI compatibility layer. [Anthropic: OpenAI SDK compatibility](https://platform.claude.com/docs/en/cli-sdks-libraries/libraries/openai-sdk)

DeepSeek follows a similar approach:

```
response = deepseek_client.responses.create(
    model="deepseek-flash",
    input="Explain MVCC.",
    metadata={
        "request_id": "abc-123"
    },
    service_tier="priority",
)
```

`metadata` and `service_tier` are unsupported, and DeepSeek states that unsupported Responses parameters are silently ignored so existing clients can connect without modification. [DeepSeek: Using the Responses API](https://api-docs.deepseek.com/guides/responses_api/)

So this:

```
assert response is not None
```

proves that the provider accepted the request. It does not prove that every requested behavior was applied.

`llm-api-adapter` keeps verified model-specific compatibility rules in its registry.

For example, OpenAI documents that when GPT-6 reasoning effort is above `none`, sampling parameters including `temperature` and `top_p` must be removed. [OpenAI: Using GPT-6](https://developers.openai.com/api/docs/guides/latest-model)

Application code can still contain:

```
response = adapter.chat(
    messages=[
        UserMessage(
            "Compare optimistic and pessimistic locking."
        )
    ],
    reasoning_level="high",
    temperature=0.2,
)
```

For `gpt-6-sol`, the registered request rule removes `temperature` before the request is sent.

When a compatibility rule removes a parameter, the adapter warns if its value differs from the registered default. In this example, `temperature=0.2` triggers a warning.

The implementation lives in the adapter's model request-rule layer: [request_rules.py](https://github.com/Inozem/llm_api_adapter/blob/main/src/llm_api_adapter/llm_registry/request_rules.py).

The goal is not to manufacture feature parity. It is to make known incompatibilities predictable instead of forcing application code to infer them from successful HTTP requests.

You can solve every example above directly:

```
if provider == "anthropic":
    ...

elif provider == "google":
    ...

elif provider == "qwen":
    ...
```

But once application code knows which provider ignores a field, which one maps it, which one uses a different reasoning scale, and which model requires parameters to be transformed or omitted, it already contains a provider abstraction — just distributed through conditionals.

A dedicated layer moves that knowledge behind one boundary:

```
application
    ↓
portable application contract
    ↓
model capability rules
    ↓
provider-specific translation
    ↓
provider API
```

For `llm-api-adapter`, the shared layer includes typed messages and responses, reasoning intent, streaming behavior, usage/cost information, and normalized error categories. Provider implementations own wire formats, request rules, reasoning mappings, response parsing, and stream-event differences.

The goal is not to make providers look identical. It is to keep differences out of application code where a stable mapping actually exists.

After that provider-specific work has been moved behind the boundary, the application call is much smaller:

```
from llm_api_adapter.models.messages.chat_message import (
    Prompt,
    UserMessage,
)
from llm_api_adapter.universal_adapter import (
    UniversalLLMAPIAdapter,
)

messages = [
    Prompt("Answer as a backend engineer."),
    UserMessage(
        "Explain connection pooling in two paragraphs."
    ),
]

adapter = UniversalLLMAPIAdapter(
    organization=organization,
    model=model,
    api_key=api_key,
)

response = adapter.chat(
    messages=messages,
    max_tokens=400,
    reasoning_level="medium",
    **provider_call_config,
)

print(response.content)
```

For most integrations:

```
provider_call_config = {}
```

Qwen is one example where a provider-specific call parameter remains necessary:

```
provider_call_config = {
    "workspace_id": QWEN_WORKSPACE_ID,
}
```

The current Qwen package uses Alibaba Model Studio's Frankfurt Global deployment through its Anthropic-compatible Messages endpoint, which requires `workspace_id` on each call. [Qwen integration README](https://github.com/Inozem/llm_api_adapter/blob/main/packages/organizations/qwen/README.md)

That requirement stays explicit, while the rest of the application call remains unchanged.

I still would not add an abstraction layer to every application.

An OpenAI-compatible endpoint can be the simpler choice when you already have an OpenAI integration, want to evaluate another model quickly, mostly need basic chat, and have tested the subset of the API your application depends on.

If your product is built deeply around one provider's unique capabilities, its native API can be the better abstraction.

A provider layer becomes more useful when **multi-provider support itself is an application requirement**. Then the important question is no longer whether another provider accepts the same JSON. It is whether the same application assumptions still hold.

**OpenAI-compatible APIs give you protocol portability. A provider abstraction can give you application-contract portability.**

Those layers overlap, but they are not the same thing.

The examples above use `llm-api-adapter` **0.9.8**. The project is open source and currently supports OpenAI, Anthropic, Google, Mistral, xAI, Qwen, Kimi, DeepSeek, and Z.ai through a mix of native and compatibility APIs.
