Deploy new inference
Choose a model and objective. InferCrane plans and runs the deployment.
Model → plan → durable deployment → endpoint
Open-source production inference
Deploy a model or connect what already runs. InferCrane operates it behind one stable endpoint.
Apache-2.0 · runs in your infrastructure · InferCrane Cloud in private preview
Three ways to begin
Choose a model and objective. InferCrane plans and runs the deployment.
Model → plan → durable deployment → endpoint
Observe vLLM, SGLang, LiteLLM, or another endpoint before changing traffic.
Observe → route → manage · no forced migration
Keep provider credentials server-side and enforce privacy and spending limits.
Stable identity · controlled spend · portable future
Cost and performance without guesswork
Compare model APIs and self-hosted plans using measured latency, throughput, reliability, and sourced cost.
Reach first traffic quickly through an approved OpenAI-compatible provider.
Provider usage · latency · errors · budget
Benchmark an exact artifact, runtime, GPU, provider, and workload shape.
TTFT · TPOT · throughput · errors · sourced cost
Promote owned capacity, retain governed fallback, or keep the API binding.
Release decision · hard budget · audit trail
Follow the operating loop from deployment through overload protection, request inspection, candidate evaluation, and recovery. Screens use representative product data.
Choose a model or bring your own. InferCrane persists intent, provisions capacity, checks readiness, and publishes the endpoint as one durable operation.
✓ Reattachable lifecycle
One product across the inference lifecycle
InferCrane owns deployment, policy, measurement, and release decisions. Runtimes, gateways, and compute remain replaceable.
Works with your inference stack
Before you try it
Yes. Start in observe-only mode with vLLM, SGLang, LiteLLM, or another OpenAI-compatible endpoint. Add traffic or lifecycle ownership only when your team is ready.
No. InferCrane operates around proven runtimes, gateways, and infrastructure. Those execution layers remain replaceable while your application keeps one model identity.
Self-hosted workloads run in infrastructure you control or behind external endpoints you connect. Provider credentials stay on the control-plane side, never in the browser.
The Apache-2.0 control plane, CLI, gateway, local quickstart, SDK source, and provider integrations are public today. The hosted web console remains a private preview.
InferCrane Cloud private preview
The open-source inference operations platform is available now. Join the waitlist for InferCrane Cloud and hands-on launch support.