cd /news/ai-infrastructure/infercrane-deploy-and-safely-evolve-… · home topics ai-infrastructure article
[ARTICLE · art-113311] src=github.com ↗ pub= topic=ai-infrastructure verified=true sentiment=· neutral

InferCrane – Deploy and safely evolve self-hosted AI inference

InferCrane, an open-source inference operations platform, released v1.0.0-rc.1, a public beta CLI and SDK that deploys, operates, and safely evolves self-hosted AI inference behind a stable OpenAI-compatible endpoint. The platform adds evidence-gated release lifecycle management—deterministic promotion, rejection, and rollback with persisted records—to raw vLLM, SGLang, custom OCI, and existing endpoints across AWS, GCP, Kubernetes, and RunPod, while preserving the active revision when evidence is missing.

read7 min views2 publishedAug 27, 2026
InferCrane – Deploy and safely evolve self-hosted AI inference
Image: Michielbdejong (auto-discovered)

Give InferCrane a model. Get a production endpoint.

Open-source infrastructure for deploying, operating, optimizing, and safely evolving

open-weight and custom-model inference behind one stable OpenAI-compatible endpoint.

Run the five-minute local proof · Read the quickstart

Plan before mutation · survive disconnects · explain latency · guard every release

Install the v1.0.0-rc.1

public beta CLI with Homebrew or use the matching SDK prerelease:

brew install infercrane/tap/infercrane
python -m pip install 'infercrane==1.0.0rc1'
npm install '@infercrane/sdk@1.0.0-rc.1'

Release archives and Terraform provider binaries are available from the v1.0.0-rc.1 prerelease. To run the complete GPU-free product proof without creating cloud resources:

git clone https://github.com/infercrane/infercrane.git
cd infercrane
make demo

The proof connects an OpenAI-compatible worker, sends and inspects a request, creates an isolated candidate, records a deterministic Release Guard rejection, verifies that production traffic did not move, and removes its disposable stack.

InferCrane is the open-source inference operations platform for self-hosted models. Your application keeps one model identity while InferCrane changes the model artifact, runtime, accelerator, provider, replica count, and active revision underneath it.

Deploy or adopt: operate vLLM, SGLang, custom OCI, and compatible existing endpoints across AWS, GCP, Kubernetes, and RunPod.Keep one endpoint: route revisions and providers behind a stable OpenAI-compatible contract.Prove every change: persist request, benchmark, quality, cost, and rollout evidence before promotion—and preserve the active revision when evidence is missing.

Raw vLLM and Kubernetes: the serving engine and manifests do not by themselves provide an evidence-gated release lifecycle. InferCrane adds deterministic promotion, rejection, rollback, and a persisted record of what changed.LiteLLM: it is an excellent routing layer. InferCrane additionally owns deployment lifecycle, evidence-gated promotion, and rollback; it can also connect to an existing LiteLLM endpoint without taking infrastructure ownership.A managed inference platform: it is often the right choice when a team wants the provider to operate its infrastructure. InferCrane is for teams that want the control plane, capacity, and billing boundary to remain in infrastructure they own.Scripts and CI: scripts can deploy a revision. InferCrane standardizes the durable reject/promote/rollback decision and the evidence attached to it.

See the detailed comparison, including the boundaries InferCrane does not own.

The local proof needs no GPU or cloud account. Real-provider support remains exact-tuple qualified; InferCrane reports missing model/runtime/hardware evidence as unknown instead of turning it into a compatibility claim.

See the system boundary

Applications and agents
          │
          ▼
Stable OpenAI-compatible endpoint
          │
          ▼
InferCrane: route · observe · optimize · release · recover
          │
          ├── deploy new inference
          ├── adopt an existing workload
          └── govern a model API or gateway
          │
          ▼
vLLM · SGLang · custom OCI · experimental Dynamo
AWS · GCP · Kubernetes · RunPod · existing infrastructure
Goal InferCrane workflow
Put a model into production Initialize a workload, review the serving plan, deploy, then call its stable endpoint.
Adopt existing inference Connect vLLM, SGLang, LiteLLM, or another compatible endpoint without transferring lifecycle ownership.
Ship a safer revision Benchmark and replay an isolated candidate, attach quality evidence, then let Release Guard promote or reject it.
Understand production failures Trace queue wait, attempts, runtime, revision, latency, saturation, and durable operations without storing prompt content.
Optimize performance and cost Propose serving configurations, measure comparable candidates on real hardware, and persist only qualified evidence.
Survive infrastructure delays Submit idempotent durable operations that continue after the CLI or control plane process disconnects.

Local fixtures prove application and lifecycle behavior. They do not prove GPU performance, cloud capacity, runtime compatibility, or model quality. InferCrane keeps those evidence boundaries explicit.

Start with a curated recipe:

infercrane workload init ./support --recipe qwen3-8b
cd support
infercrane workload plan
infercrane workload deploy --wait

Or bring another compatible immutable model identity:

infercrane workload init ./agent-model --model mistralai/Mistral-7B-Instruct-v0.3
cd agent-model
infercrane workload plan
infercrane workload deploy --wait

Recipes are reproducible configuration starting points, not benchmark claims or an allowlist. Evidence remains bound to the exact model commit, runtime, accelerator, provider, cache state, and workload.

Already operating a workload? Connect it first:

infercrane connect https://vllm.internal/v1 --as support-production
infercrane doctor support-production
infercrane observe support-production

InferCrane can observe an existing endpoint before it manages traffic or infrastructure. Provider credentials and request content do not enter the browser console.

InferCrane separates a modeled proposal from measured and qualified evidence:

infercrane optimize propose llama-3.1-8b-instruct \
  --provider aws \
  --region eu-central-1 \
  --gpu L40S \
  --objective interactive \
  --write-dir .infercrane/candidates
model + hardware + workload + SLO + cost target
                       │
                       ▼
              candidate serving plans
                       │
                       ▼
          AIPerf + replay + quality evidence
                       │
                       ▼
            performance · errors · cost
                       │
                 ┌─────┴─────┐
                 ▼           ▼
              promote      reject

InferCrane composes replaceable execution technology instead of rebuilding it. vLLM, SGLang, and custom OCI are current execution paths. Dynamo is experimental. TensorRT-LLM, LMCache, NIXL, and external optimizers remain capability boundaries until their exact adapters and hardware tuples are qualified. InferCrane owns serving-plan identity, durable operations, comparable evidence, routing policy, promotion, and rollback.

Stable endpoint identity: applications do not change when the serving plan changes.Bounded overload: admission limits, explicit429

andRetry-After

, one end-to-end deadline, and bounded retries prevent unlimited queue growth.Request-path isolation: gateways route from immutable in-memory snapshots and never query PostgreSQL on the inference request path.Durable operations: deployment, scaling, deletion, and release work is idempotent, restart-safe, cancellable, and inspectable.Release evidence: benchmark, replay, quality, reliability, and sourced cost evidence can block a candidate before traffic moves.Content-free operations: request evidence records operational metadata without persisting prompts or model outputs.Explicit ownership: existing runtimes, gateways, training systems, sandboxes, and clouds stay replaceable behind versioned contracts.

Read the architecture, system invariants, and data flows for the complete design.

Interface Status and purpose
CLI and control API Primary deployment, operation, evidence, and administration interfaces.
OpenAI-compatible gateway Capability-gated Chat, Completions, Embeddings, Responses, and online batch paths. The pinned vLLM profile currently qualifies Chat plus model-compatible Completions and Embeddings; unsupported capabilities fail before upstream transmission.
Python and TypeScript SDKs Public beta packages: infercrane==1.0.0rc1 and @infercrane/sdk@1.0.0-rc.1 . Generated from the checked OpenAPI contract.
Terraform provider Logical deployment lifecycle with guarded updates and import. Release binaries and source are public; Registry publication is pending.
Terminal workspace Fleet attention, evidence inspection, and state-valid guarded actions.
Browser console Separate deny-by-default private-preview application using the same control API.
Read-only MCP server Closed-world operational inspection without deployment, scaling, promotion, deletion, budget, or secret tools.

InferCrane v1.0.0-rc.1

is the first public beta. The stable v1.0.0

release will promote the exact qualified product contract after the prerelease cycle; no earlier development tag should be treated as a supported public release.

  • Local race, PostgreSQL, fault-injection, Docker, Kind, KWOK, package, migration, security, and documentation gates are automated.
  • AWS has exact-tuple real GPU evidence for vLLM, SGLang, custom OCI, model identity, requests, bounded benchmarks, durable deletion, and final zero managed-resource inventory.
  • GCP GPU, real GPU Kubernetes/KServe, additional model/runtime/GPU tuples, and several distributed optimization paths still require separate real-infrastructure evidence.
  • No benchmark is generalized beyond the exact tuple and workload that produced it.

See the authoritative compatibility and qualification policy, AWS evidence, and feature qualification matrix before relying on an exact provider, runtime, model, or accelerator combination.

Five-minute quickstartProduct conceptsBuild new inferenceConnect existing inferenceSafe releasesProvider setupProduction operationsPython SDKTypeScript SDKTerraform providerSecuritySupport

Mintlify generates llms.txt and

from the public documentation. Every public documentation page is also available as Markdown by appending

llms-full.txt

.md

to its URL.Contributions are welcome. Start with CONTRIBUTING.md, follow the Code of Conduct, and sign commits with git commit -s

. Changes must include tests and relevant documentation. Durable architecture, security, storage, and dependency changes must update their authoritative public documentation.

Never disclose credentials, prompts, model responses, private endpoints, or suspected vulnerabilities in a public issue. Use the private reporting process in SECURITY.md. Questions and reproducible defects follow SUPPORT.md.

InferCrane Community is available under the Apache License 2.0. Hosted and enterprise products are separate distributions and are not licensed by this repository. Release archives also include third-party notices and a release-specific SPDX SBOM. The InferCrane name and crane logo remain subject to the trademark policy.

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @infercrane 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/infercrane-deploy-an…] indexed:0 read:7min 2026-08-27 ·