{"slug": "show-hn-infercrane-one-stable-endpoint-for-self-hosted-ai-inference", "title": "Show HN: InferCrane – One stable endpoint for self-hosted AI inference", "summary": "InferCrane, an open-source inference operations platform, launched with a stable endpoint for self-hosted AI inference, supporting vLLM, SGLang, LiteLLM, and other OpenAI-compatible endpoints. The Apache-2.0 control plane, CLI, gateway, local quickstart, SDK source, and provider integrations are public today, with InferCrane Cloud in private preview. The platform offers observe-only mode, cost and performance benchmarking, and durable deployment lifecycle management.", "body_md": "### Deploy new inference\n\nChoose a model and objective. InferCrane plans and runs the deployment.\n\n`Model → plan → durable deployment → endpoint`\n\nOpen-source production inference\n\nDeploy a model or connect what already runs. InferCrane operates it behind one stable endpoint.\n\nApache-2.0 · runs in your infrastructure · InferCrane Cloud in private preview\n\nThree ways to begin\n\nChoose a model and objective. InferCrane plans and runs the deployment.\n\n`Model → plan → durable deployment → endpoint`\n\nObserve vLLM, SGLang, LiteLLM, or another endpoint before changing traffic.\n\n`Observe → route → manage · no forced migration`\n\nKeep provider credentials server-side and enforce privacy and spending limits.\n\n`Stable identity · controlled spend · portable future`\n\nCost and performance without guesswork\n\nCompare model APIs and self-hosted plans using measured latency, throughput, reliability, and sourced cost.\n\nReach first traffic quickly through an approved OpenAI-compatible provider.\n\n`Provider usage · latency · errors · budget`\n\nBenchmark an exact artifact, runtime, GPU, provider, and workload shape.\n\n`TTFT · TPOT · throughput · errors · sourced cost`\n\nPromote owned capacity, retain governed fallback, or keep the API binding.\n\n`Release decision · hard budget · audit trail`\n\nFollow the operating loop from deployment through overload protection, request inspection, candidate evaluation, and recovery. Screens use representative product data.\n\nChoose a model or bring your own. InferCrane persists intent, provisions capacity, checks readiness, and publishes the endpoint as one durable operation.\n\n`✓ Reattachable lifecycle`\n\nOne product across the inference lifecycle\n\nInferCrane owns deployment, policy, measurement, and release decisions. Runtimes, gateways, and compute remain replaceable.\n\nWorks with your inference stack\n\nBefore you try it\n\nYes. Start in observe-only mode with vLLM, SGLang, LiteLLM, or another OpenAI-compatible endpoint. Add traffic or lifecycle ownership only when your team is ready.\n\nNo. InferCrane operates around proven runtimes, gateways, and infrastructure. Those execution layers remain replaceable while your application keeps one model identity.\n\nSelf-hosted workloads run in infrastructure you control or behind external endpoints you connect. Provider credentials stay on the control-plane side, never in the browser.\n\nThe Apache-2.0 control plane, CLI, gateway, local quickstart, SDK source, and provider integrations are public today. The hosted web console remains a private preview.\n\nInferCrane Cloud private preview\n\nThe open-source inference operations platform is available now. Join the waitlist for InferCrane Cloud and hands-on launch support.", "url": "https://wpnews.pro/news/show-hn-infercrane-one-stable-endpoint-for-self-hosted-ai-inference", "canonical_source": "https://infercrane.com", "published_at": "2026-08-28 14:34:11+00:00", "updated_at": "2026-08-28 14:48:25.695235+00:00", "lang": "en", "topics": ["ai-infrastructure", "ai-tools", "mlops", "developer-tools"], "entities": ["InferCrane", "vLLM", "SGLang", "LiteLLM"], "alternates": {"html": "https://wpnews.pro/news/show-hn-infercrane-one-stable-endpoint-for-self-hosted-ai-inference", "markdown": "https://wpnews.pro/news/show-hn-infercrane-one-stable-endpoint-for-self-hosted-ai-inference.md", "text": "https://wpnews.pro/news/show-hn-infercrane-one-stable-endpoint-for-self-hosted-ai-inference.txt", "jsonld": "https://wpnews.pro/news/show-hn-infercrane-one-stable-endpoint-for-self-hosted-ai-inference.jsonld"}}