# Show HN: InferCrane – One stable endpoint for self-hosted AI inference

> Source: <https://infercrane.com>
> Published: 2026-08-28 14:34:11+00:00

### Deploy new inference

Choose a model and objective. InferCrane plans and runs the deployment.

`Model → plan → durable deployment → endpoint`

Open-source production inference

Deploy a model or connect what already runs. InferCrane operates it behind one stable endpoint.

Apache-2.0 · runs in your infrastructure · InferCrane Cloud in private preview

Three ways to begin

Choose a model and objective. InferCrane plans and runs the deployment.

`Model → plan → durable deployment → endpoint`

Observe vLLM, SGLang, LiteLLM, or another endpoint before changing traffic.

`Observe → route → manage · no forced migration`

Keep provider credentials server-side and enforce privacy and spending limits.

`Stable identity · controlled spend · portable future`

Cost and performance without guesswork

Compare model APIs and self-hosted plans using measured latency, throughput, reliability, and sourced cost.

Reach first traffic quickly through an approved OpenAI-compatible provider.

`Provider usage · latency · errors · budget`

Benchmark an exact artifact, runtime, GPU, provider, and workload shape.

`TTFT · TPOT · throughput · errors · sourced cost`

Promote owned capacity, retain governed fallback, or keep the API binding.

`Release decision · hard budget · audit trail`

Follow the operating loop from deployment through overload protection, request inspection, candidate evaluation, and recovery. Screens use representative product data.

Choose a model or bring your own. InferCrane persists intent, provisions capacity, checks readiness, and publishes the endpoint as one durable operation.

`✓ Reattachable lifecycle`

One product across the inference lifecycle

InferCrane owns deployment, policy, measurement, and release decisions. Runtimes, gateways, and compute remain replaceable.

Works with your inference stack

Before you try it

Yes. Start in observe-only mode with vLLM, SGLang, LiteLLM, or another OpenAI-compatible endpoint. Add traffic or lifecycle ownership only when your team is ready.

No. InferCrane operates around proven runtimes, gateways, and infrastructure. Those execution layers remain replaceable while your application keeps one model identity.

Self-hosted workloads run in infrastructure you control or behind external endpoints you connect. Provider credentials stay on the control-plane side, never in the browser.

The Apache-2.0 control plane, CLI, gateway, local quickstart, SDK source, and provider integrations are public today. The hosted web console remains a private preview.

InferCrane Cloud private preview

The open-source inference operations platform is available now. Join the waitlist for InferCrane Cloud and hands-on launch support.
