cd /news/ai-tools/openai-compatible-apis-vs-provider-a… · home › topics › ai-tools › article
[ARTICLE · art-147721] src=dev.to ↗ pub= topic=ai-tools verified=true sentiment=· neutral

OpenAI-compatible APIs vs provider abstraction: What actually transfers across LLM providers?

A developer built llm-api-adapter, an open-source Python SDK that provides a provider-neutral contract across OpenAI, Anthropic, Google, Mistral, xAI, Qwen, Kimi, DeepSeek, and Z.ai, to address portability gaps between OpenAI-compatible APIs. The writeup documents that compatibility operates on a ladder from transport to operational levels, and that Anthropic's OpenAI compatibility endpoint hoists and concatenates system and developer messages into a single initial system message, so instruction placement and role distinctions are not preserved even though the request succeeds.

by read8 min views2 publishedOct 8, 2026

OpenAI-compatible APIs promise something very attractive: write one integration, change the base_url and model, and move between LLM providers without rewriting your application.

For basic chat, this often works surprisingly well. The more interesting question is what that compatibility actually guarantees.

The HTTP format? The accepted parameters? Or the behavior your application can rely on?

I ran into this distinction repeatedly while building llm-api-adapter, an open-source, lightweight Python SDK that provides a provider-neutral contract across OpenAI, Anthropic, Google, Mistral, xAI, Qwen, Kimi, DeepSeek, and Z.ai.

I’ll use it throughout this article as one example of how these portability problems can be handled. The point is not that every application needs another abstraction layer. In many cases, an OpenAI-compatible endpoint is enough.

The goal is to find where it stops being enough.

For this comparison, I’m staying mostly inside the chat contract: instruction semantics, reasoning controls, and unsupported or model-specific parameters. Tool calling and structured outputs introduce another set of differences and deserve separate comparisons.

Verification note: Provider documentation was checked on October 8, 2026. Library examples use llm-api-adapter 0.9.8 and follow its tested public API.

Each provider example assumes an OpenAI client configured with that provider’s API key and base URL. I use names such as anthropic_client, gemini_client, and qwen_client below to make that explicit.

At the simplest level, compatibility looks like this:

from openai import OpenAI

client = OpenAI(
    api_key=PROVIDER_API_KEY,
    base_url=PROVIDER_BASE_URL,
)

response = client.chat.completions.create(
    model=MODEL,
    messages=[
        {
            "role": "user",
            "content": "Why do database indexes speed up reads?"
        }
    ],
)

print(response.choices[0].message.content)

For a simple user-message → text-response workflow, that may be all you need.

But production applications tend to depend on more: instruction roles, reasoning controls, usage metadata, and failure behavior.

I find it useful to think about compatibility as a ladder:

transport compatibility
        ↓
request-shape compatibility
        ↓
feature compatibility
        ↓
semantic compatibility
        ↓
operational compatibility

Two APIs can be compatible near the top and still require different application logic further down.

Consider an OpenAI-style conversation:

messages = [
    {
        "role": "system",
        "content": "Follow company policy A."
    },
    {
        "role": "user",
        "content": "First question"
    },
    {
        "role": "developer",
        "content": "For the next step, prefer policy B."
    },
    {
        "role": "user",
        "content": "Second question"
    },
]

Anthropic's OpenAI compatibility endpoint accepts that shape:

response = anthropic_client.chat.completions.create(
    model="claude-opus-5-5",
    messages=messages,
)

But the compatibility layer documents a transformation: system and developer messages are hoisted to the beginning and concatenated into one initial system message. It also documents a number of OpenAI parameters that are accepted but ignored. Anthropic: OpenAI SDK compatibility

Conceptually, this:

system: Follow company policy A.

user: First question

developer: For the next step, prefer policy B.

user: Second question

is represented more like:

system:
Follow company policy A.
For the next step, prefer policy B.

user: First question

user: Second question

The request succeeds, but instruction placement and role distinctions are not preserved.

This is specifically a property of Anthropic's OpenAI compatibility layer. The native Claude API supports mid-conversation system messages on selected models, including Opus 5.5, subject to documented placement constraints. Anthropic: Mid-conversation system messages

If the requirement can be expressed as a portable initial instruction, the application can use Prompt:

from llm_api_adapter.models.messages.chat_message import (
    Prompt,
    UserMessage,
)

messages = [
    Prompt("Follow company policy A."),
    UserMessage("First question"),
    UserMessage("Second question"),
]

response = adapter.chat(messages=messages)

If the application specifically depends on a mid-conversation OpenAI developer role, that distinction is not treated as universally portable. It should remain provider-specific instead of being silently translated into something that means something else.

Sometimes the correct normalization is to expose a boundary rather than hide it.

reasoning_effort parameter can mean several different things Reasoning controls are a useful example because several providers deliberately expose OpenAI-style parameters.

Suppose the application asks for:

reasoning_effort="high"

That looks portable. Its meaning is not.

Anthropic's compatibility endpoint accepts:

response = anthropic_client.chat.completions.create(
    model="claude-opus-5-5",
    messages=messages,
    reasoning_effort="high",
)

But Anthropic documents reasoning_effort as ignored by this compatibility layer. Claude has native mechanisms for controlling reasoning behavior, and Anthropic recommends the native API when applications need access to the full thinking feature set. Anthropic: OpenAI SDK compatibility

Google takes a different approach:

response = gemini_client.chat.completions.create(
    model="gemini-3.6-flash",
    messages=messages,
    reasoning_effort="high",
)

This example uses the stable gemini-3.6-flash model. Google's OpenAI compatibility layer maps reasoning_effort onto Gemini's own thinking controls rather than simply ignoring it. Google: OpenAI compatibility

Qwen 3.8 adds another interpretation:

response = qwen_client.chat.completions.create(
    model="qwen3.8-flash",
    messages=messages,
    reasoning_effort="high",
)

Qwen 3.8 uses low, medium, and xhigh as its native reasoning levels, with xhigh as the default. OpenAI-compatible values are mapped onto that scale: for example, minimal maps to low, while high and max map to xhigh. none disables thinking.

Invalid values can cause an error, and reasoning_effort cannot be supplied together with thinking_budget. Qwen: OpenAI-compatible Chat API

So the same-looking setting can mean:

Claude
high → ignored by compatibility layer

Gemini
high → translated to Gemini thinking controls

Qwen
high → mapped to xhigh

The application uses one setting:

response = adapter.chat(
    messages=[
        UserMessage("Solve this architecture problem.")
    ],
    reasoning_level="high",
)

reasoning_level expresses configuration intent. For categorical models, the adapter projects portable levels onto the model's registered native levels. For models with numeric reasoning budgets, it computes corresponding values within the registered supported range.

Portable levels are:

none
minimal
low
medium
high
very_high

Provider-native values such as xhigh can still be used for models that support them, but they are intentionally not considered portable.

These mappings normalize configuration intent; they do not guarantee equivalent reasoning depth, quality, latency, or cost across models.

The abstraction can normalize how the application asks for reasoning. It cannot make different models reason equivalently.

Compatibility layers often ignore unsupported fields instead of failing the request. That is useful for migrations, but it changes what a successful response proves.

For Anthropic:

response = anthropic_client.chat.completions.create(
    model="claude-opus-5-5",
    messages=messages,
    reasoning_effort="high",
    service_tier="priority",
)

Anthropic currently documents both fields as ignored through its OpenAI compatibility layer. Anthropic: OpenAI SDK compatibility

DeepSeek follows a similar approach:

response = deepseek_client.responses.create(
    model="deepseek-flash",
    input="Explain MVCC.",
    metadata={
        "request_id": "abc-123"
    },
    service_tier="priority",
)

metadata and service_tier are unsupported, and DeepSeek states that unsupported Responses parameters are silently ignored so existing clients can connect without modification. DeepSeek: Using the Responses API

So this:

assert response is not None

proves that the provider accepted the request. It does not prove that every requested behavior was applied.

llm-api-adapter keeps verified model-specific compatibility rules in its registry.

For example, OpenAI documents that when GPT-6 reasoning effort is above none, sampling parameters including temperature and top_p must be removed. OpenAI: Using GPT-6

Application code can still contain:

response = adapter.chat(
    messages=[
        UserMessage(
            "Compare optimistic and pessimistic locking."
        )
    ],
    reasoning_level="high",
    temperature=0.2,
)

For gpt-6-sol, the registered request rule removes temperature before the request is sent.

When a compatibility rule removes a parameter, the adapter warns if its value differs from the registered default. In this example, temperature=0.2 triggers a warning.

The implementation lives in the adapter's model request-rule layer: request_rules.py.

The goal is not to manufacture feature parity. It is to make known incompatibilities predictable instead of forcing application code to infer them from successful HTTP requests.

You can solve every example above directly:

if provider == "anthropic":
    ...

elif provider == "google":
    ...

elif provider == "qwen":
    ...

But once application code knows which provider ignores a field, which one maps it, which one uses a different reasoning scale, and which model requires parameters to be transformed or omitted, it already contains a provider abstraction — just distributed through conditionals.

A dedicated layer moves that knowledge behind one boundary:

application
    ↓
portable application contract
    ↓
model capability rules
    ↓
provider-specific translation
    ↓
provider API

For llm-api-adapter, the shared layer includes typed messages and responses, reasoning intent, streaming behavior, usage/cost information, and normalized error categories. Provider implementations own wire formats, request rules, reasoning mappings, response parsing, and stream-event differences.

The goal is not to make providers look identical. It is to keep differences out of application code where a stable mapping actually exists.

After that provider-specific work has been moved behind the boundary, the application call is much smaller:

from llm_api_adapter.models.messages.chat_message import (
    Prompt,
    UserMessage,
)
from llm_api_adapter.universal_adapter import (
    UniversalLLMAPIAdapter,
)

messages = [
    Prompt("Answer as a backend engineer."),
    UserMessage(
        "Explain connection pooling in two paragraphs."
    ),
]

adapter = UniversalLLMAPIAdapter(
    organization=organization,
    model=model,
    api_key=api_key,
)

response = adapter.chat(
    messages=messages,
    max_tokens=400,
    reasoning_level="medium",
    **provider_call_config,
)

print(response.content)

For most integrations:

provider_call_config = {}

Qwen is one example where a provider-specific call parameter remains necessary:

provider_call_config = {
    "workspace_id": QWEN_WORKSPACE_ID,
}

The current Qwen package uses Alibaba Model Studio's Frankfurt Global deployment through its Anthropic-compatible Messages endpoint, which requires workspace_id on each call. Qwen integration README

That requirement stays explicit, while the rest of the application call remains unchanged.

I still would not add an abstraction layer to every application.

An OpenAI-compatible endpoint can be the simpler choice when you already have an OpenAI integration, want to evaluate another model quickly, mostly need basic chat, and have tested the subset of the API your application depends on.

If your product is built deeply around one provider's unique capabilities, its native API can be the better abstraction.

A provider layer becomes more useful when multi-provider support itself is an application requirement. Then the important question is no longer whether another provider accepts the same JSON. It is whether the same application assumptions still hold.

OpenAI-compatible APIs give you protocol portability. A provider abstraction can give you application-contract portability.

Those layers overlap, but they are not the same thing.

The examples above use llm-api-adapter 0.9.8. The project is open source and currently supports OpenAI, Anthropic, Google, Mistral, xAI, Qwen, Kimi, DeepSeek, and Z.ai through a mix of native and compatibility APIs.

── more in #ai-tools 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/openai-compatible-ap…] indexed:0 read:8min 2026-10-08 · —