cd /news/ai-infrastructure/aicr-v1-0-open-stable-and-verifiable… · home › topics › ai-infrastructure › article
[ARTICLE · art-146177] src=developer.nvidia.com ↗ pub= topic=ai-infrastructure verified=true sentiment=↑ positive

AICR v1.0: Open, stable, and verifiable GPU cluster configuration

NVIDIA released AICR v1.0, an open-source AI Cluster Runtime that provides version-locked, validated recipes for GPU-accelerated Kubernetes cluster configuration, establishing a stable compatibility contract across its CLI, REST API, Go SDK, bundle layout, and artifact schemas. The project now has over 100 distinct contributors, almost half from outside NVIDIA, and is integrated by Pulumi Labs as an infrastructure-as-code provider and by Mirantis's k0rdent for multi-cluster management. Each recipe pins component combinations that work together, renders deployment artifacts for Helm, Argo CD, Flux, or Helmfile, and carries signed validation evidence from the hardware it was tested on, with a validation dashboard at validation.aicr.run for finding recipes by service, GPU, operating system, workload intent, and platform.

by read5 min views3 publishedOct 6, 2026
AICR v1.0: Open, stable, and verifiable GPU cluster configuration
Image: NVIDIA Developer Blog

GPU-accelerated Kubernetes clusters depend on compatible versions across dozens of components, each on its own release cycle: host kernels, GPU drivers, container runtimes, networking, storage, operators, and workload frameworks.

A configuration that works for one service, GPU generation, and Kubernetes release may silently fail for another, and tracing version conflicts after deployment is slow and error-prone.

NVIDIA AI Cluster Runtime (AICR) addresses this with version-locked, validated recipes for GPU cluster configuration. Each recipe pins the component combinations that work together, renders deployment artifacts for Helm, Argo CD, Flux, or Helmfile, and carries signed validation evidence from the hardware it was tested on.

The v1.0 release of AICR establishes a stable compatibility contract across its CLI, REST API, Go SDK, bundle layout, and artifact schemas so operators, integrators, and contributors can build on AICR’s public interfaces with confidence.

The validation dashboard lets operators find recipes by service, GPU, operating system, workload intent, and optional platform, then inspect each recipe’s status and any published evidence for the hardware configuration tested. Integrators can build against AICR’s public interfaces under the v1.x compatibility policy. Contributors can propose recipes for environments the maintainers cannot test, validate them on their own clusters, and submit signed evidence for maintainer review.

That recipe model is also finding uses across the ecosystem. Pulumi Labs exposes AICR through an infrastructure-as-code provider, while Mirantis’s k0rdent integration packages it for multi-cluster management. Together, they demonstrate the value of defining GPU-accelerated Kubernetes configuration once and consuming it through different tools. Today, AICR has over 100 distinct contributors, with almost half from outside of NVIDIA!

Why GPU cluster configuration needs a reproducible contract #

GPU-accelerated Kubernetes clusters depend on compatible versions and settings across host kernels, GPU drivers, container runtimes, Kubernetes, networking, storage, device plugins, operators, schedulers, and workload frameworks. These components follow different release cycles; upgrading one can break a previously working combination. A configuration validated for one service, GPU generation, fabric type, machine shape, and Kubernetes release may fail for another, and small version differences can be difficult to trace after deployment.

Even if every component installs successfully, the cluster may not meet the recipe’s intended configuration. Installation doesn’t confirm that components are healthy, that required capabilities like gang scheduling or accelerator discovery work, or that measured results meet a recipe’s performance thresholds.

Knowledge of which combinations work and how they were validated has lived in separate validation systems, deployment scripts, and operational runbooks. That makes it hard for teams to discover, reproduce, and update working configurations.

Over the past six months, AICR has grown from a handful of recipes to a library spanning major Kubernetes services and the current NVIDIA accelerator portfolio, rendered as deployer-neutral bundles. We added live-cluster validation, signed evidence, public evidence aggregation, and supply-chain verification. For v1.0, we also added committed compatibility baselines and merge-blocking checks around the public integration surfaces.

From observed state to a verifiable result #

AICR provides four core capabilities:

  • Snapshot records observed cluster state, including Kubernetes, operating system, kernel, GPU, and topology information.
  • Recipe describes the desired, version-locked component configuration and the constraints and validation phases that apply to it.
  • Bundle renders the recipe into artifacts for the operator’s preferred deployment tooling.
  • Validation compares the recipe with observed state and, where declared, runs deployment, conformance, and performance checks against the cluster.

These four capabilities are deliberately independent. A snapshot records the observed state; it is not a desired configuration. A recipe describes the desired configuration; it does not reconcile a cluster. Common open source CD tools like Helm, Argo CD, Flux, or Helmfile apply or reconcile the bundle into the cluster. AICR can then validate the running cluster against the recipe, and record signed evidence of the result.

These capabilities can be combined in multiple sequences. Snapshot data or explicit target criteria can produce a recipe. A recipe can produce a bundle for an existing deployer. The recipe and observed cluster state feed validation. An operator explicitly verifies bundles and evidence when they invoke the corresponding command.

For example, an operator can select EKS, GB300, Ubuntu, training, and Kubeflow; resolve those criteria to a pinned recipe; render the recipe for Argo CD; deploy it through the existing GitOps workflow; and validate the running cluster against the same recipe. The intended configuration does not change if the operator instead renders it for Helm, Flux, or Helmfile.

What’s new in AICR v1.0 #

AICR v1.0 defines compatibility rules for its public CLI, REST API, Go SDK, bundle layout, and artifact schemas. It also lets operators inspect the validation evidence published for each Supported recipe: what was tested, which checks passed, and who signed the results.

AICR v1.0 sets compatibility rules for:

  • aicr CLI’s public commands, flags, exit semantics, and structured output
  • aicrd REST API and OpenAPI contract
  • exported API of the github.com/NVIDIA/aicr/pkg/client/v1 package
  • generated bundle layout and AICR artifact schemas

Each public interface has a committed baseline checked before changes merge. The release policy also defines semantic breaking changes. After v1.0, removing or incompatibly changing a stable public interface requires a new major release.

For Go integrators, pkg/client/v1 exposes the supported workflow without requiring imports from AICR’s internal packages. The CLI and REST server use the same facade, reducing the risk that one public entry point behaves differently from another.

Try AICR and contribute #

Try a recipe for your environment, inspect its status and any published evidence, and run the dashboard’s aicr evidence verify command where evidence is available. Contributions are especially useful for hardware and cluster combinations outside current project coverage. You can:

  • Contribute features, integrations, or documentation
  • Propose a recipe for hardware, OS, or a service not yet represented in AICR
  • Validate a recipe in your own cluster and submit signed evidence
  • Report bugs, share feedback, or request features through GitHub Issues

Start with the project repository, contributing guide, and issue tracker.

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @nvidia 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/aicr-v1-0-open-stabl…] indexed:0 read:5min 2026-10-06 · —