{"slug": "best-practices-for-amazon-sagemaker-hyperpod-administration-and-governance", "title": "Best practices for Amazon SageMaker HyperPod administration and governance", "summary": "Amazon published best practices for administering Amazon SageMaker HyperPod through Amazon SageMaker Unified Studio, laying out four layers of control — organization, project, cluster, and workload — for ML teams sharing accelerated compute. The guidance states that a SageMaker HyperPod connection adds an approved cluster to a project but does not replace the cluster's AWS Identity and Access Management (IAM), Amazon Elastic Kubernetes Service (Amazon EKS), or Slurm controls. The model is intended to let infrastructure teams keep cluster operations while offering approved compute to ML teams in their project context.", "body_md": "## [Artificial Intelligence](https://aws.amazon.com/blogs/machine-learning/)\n\n# Best practices for Amazon SageMaker HyperPod administration and governance\n\n[Amazon SageMaker HyperPod](https://aws.amazon.com/sagemaker/hyperpod/) gives machine learning (ML) teams access to large pools of accelerated compute for training and fine-tuning models. When several teams share one cluster, the technical setup is usually straightforward. The challenging part is governance. You must decide which teams can use the cluster, how much capacity each team gets, what happens when one team’s workload competes with another’s, and who is accountable when usage drifts from policy. [Amazon SageMaker Unified Studio](https://aws.amazon.com/sagemaker/unified-studio/) adds another consideration: You can connect a SageMaker HyperPod cluster to a project so team members can launch workloads from their project workspace. That convenience is valuable, but after multiple teams share visibility into the same cluster, the controls that govern who can do what become even more important. In this post, we show how to administer SageMaker HyperPod through SageMaker Unified Studio while preserving the underlying governance controls. We cover the four layers of control: organization, project, cluster, and workload. We also explain how to design identity, capacity, and observability policies across them. By the end, you will have a repeatable model for offering approved SageMaker HyperPod compute to ML teams in their project context while keeping cluster operations with the infrastructure team.\n\nAmazon SageMaker HyperPod is a capability of [Amazon SageMaker AI](https://aws.amazon.com/sagemaker/ai/). Amazon SageMaker Unified Studio is the data and AI development environment where teams build with their data and tools. With SageMaker Unified Studio, you can connect a project to an existing SageMaker HyperPod cluster. Members can then launch machine learning workloads, review cluster and task information, and open a JupyterLab workflow. You continue to manage clusters through Amazon SageMaker AI interfaces and APIs.\n\nThis separation gives infrastructure teams a useful operating model. You can manage cluster infrastructure through established cloud operations processes. At the same time, you can present approved compute to machine learning teams in their project context. In this post, we describe how you can design infrastructure boundaries, govern access, allocate shared capacity, and operate SageMaker HyperPod consistently through SageMaker Unified Studio.\n\n## Understand the administrative boundaries\n\nA well-governed environment separates organizational administration, project access, and cluster operations. Each layer answers a different question and uses a different control. The following table summarizes each boundary, its primary controls, and its administrative purpose.\n\n| **Boundary** | **Primary controls** | **Administrative purpose** | \n| Organization | SageMaker Unified Studio domains, domain units, associated accounts, project profiles, and authorization policies | Determines who can create projects, which accounts and Regions projects can use, and which tools are available | \n| Project | Project membership, project roles, and SageMaker HyperPod connections | Defines the collaboration context and the AWS resources that project members can access | \n| Cluster | SageMaker HyperPod cluster admin roles, Amazon EKS access entries, role-based access control (RBAC), and EKS Pod Identity, or Slurm controls | Governs cluster configuration, scheduler access, namespaces, tasks, and infrastructure operations | \n| Workload | Compute allocations, priority classes, lending and borrowing policies, and task permissions | Controls who can submit work and how shared capacity is assigned | \n\nTreat these controls as layers. A SageMaker HyperPod connection adds an approved cluster to a project. It doesn’t replace the cluster’s AWS Identity and Access Management (IAM), Amazon Elastic Kubernetes Service (Amazon EKS), or Slurm controls. Review the project role, connection access role, EKS access entries and RBAC or Slurm controls, workload identity, data and [AWS Key Management Service (AWS KMS) key](https://aws.amazon.com/kms/) policies, network policy, task-view restrictions, and scheduler policy together. Make the cluster available only after those controls are aligned.\n\nThese controls form four layers, shown in the following figure.\n\nSageMaker Unified Studio projects are collaboration boundaries. They aren’t strong runtime security boundaries. Keep the SageMaker HyperPod cluster, scheduler, and scarce accelerator capacity under one designated capacity account. Approved users and datasets can remain in the same account or in separate consumer and data accounts. AWS documents multi-account support for SageMaker HyperPod task governance on Amazon EKS clusters. Although On-Demand Capacity Reservations can be shared across accounts, centralize SageMaker HyperPod cluster ownership and capacity administration in the capacity account. Use approved cross-account access rather than making each consumer account an independent capacity administrator.\n\n- For Amazon EKS, use a namespace per tenant together with RBAC, per-tenant service accounts and EKS Pod Identity roles, default-deny network policies, and tenant-specific storage and AWS KMS permissions. For Slurm, use Slurm accounting with hierarchical accounts and associations, quality of service (QoS), priority and fair-share policies, and partitions. Also use operating-system identity and file permissions, network controls, and tenant-specific data paths.\n- Use IAM roles and resource policies, rather than project membership alone, to control access to Amazon Simple Storage Service (Amazon S3) buckets, AWS KMS keys, secrets, container registries, and other data services. Use project and domain-unit policies for collaboration and delegation.\n- For Amazon EKS clusters, use SageMaker HyperPod task governance, including quotas, priority classes, lending and borrowing, and preemption, to share the central accelerator pool fairly. For Slurm clusters, use native partitions, quality of service (QoS), priority, fair-share, and preemption controls. Slurm doesn’t provide the same lending-and-borrowing model. These scheduling controls determine when an authorized workload receives compute. They don’t grant namespace or data access.\n\nFor stronger isolation within a cluster, use dedicated nodes or node groups and admission controls where appropriate. Use a separate cluster or account when legal, regulatory, or security requirements demand hard infrastructure isolation. Cross-account consumer access can preserve account-level ownership without duplicating the central SageMaker HyperPod cluster. Within the capacity account, use [project profiles](https://docs.aws.amazon.com/sagemaker-unified-studio/latest/adminguide/custom.html) and [domain units](https://docs.aws.amazon.com/sagemaker-unified-studio/latest/adminguide/domain-units.html) with project authorization policies to delegate project creation and ownership.\n\nThe following figure shows centralized capacity with tenant-specific workload, identity, data, network, and visibility controls.\n\nFigure 3 shows how those boundaries connect when SageMaker Unified Studio exposes approved SageMaker HyperPod compute to project members.\n\n## Decide when SageMaker Unified Studio is the right administrative experience\n\nWith SageMaker Unified Studio, machine learning teams that already work in a project can follow an approved path to shared SageMaker HyperPod compute. Members can find connected clusters, review status and metadata, inspect supported tasks and metrics, and move into JupyterLab without using a separate infrastructure inventory.\n\nThis experience doesn’t replace cluster administration. The following table shows how each persona should use SageMaker Unified Studio alongside the service-specific interfaces.\n\n| **Persona** | **Use SageMaker Unified Studio for** | **Continue using service-specific tools for** | \n| Domain or infrastructure administrator | Domain units, project creation policy, project profiles, membership policy, and account placement | Organization-level controls, account provisioning, and infrastructure automation | \n| SageMaker HyperPod cluster administrator | Presenting approved cluster connections and reviewing cluster, task, settings, and metadata views | Cluster creation, updates, resiliency configuration, add-ons, EKS or Slurm administration, and incident response | \n| Project owner | Managing project membership and giving users a consistent project context for approved compute | Requesting infrastructure changes and approving business-specific access requirements | \n| ML engineer or data scientist | Finding approved compute, reviewing workload status, and opening the JupyterLab workflow | Submitting and managing detailed workloads through the SageMaker [HyperPod CLI](https://docs.aws.amazon.com/sagemaker/latest/dg/getting-started-hyperpod-training-deploying-models.html) ,`kubectl` , or Slurm tools as appropriate | \n\nUse SageMaker Unified Studio when project membership, data access, development tools, and compute need a common context. For cluster changes and repeatable automation, continue using SageMaker AI APIs, infrastructure as code, and orchestrator tools. For reference implementations and integrations, refer to the [AI on SageMaker HyperPod](https://awslabs.github.io/ai-on-sagemaker-hyperpod) site.\n\n## Make infrastructure decisions before connecting a cluster\n\nBefore approving a connection, document the capacity account, any consumer or data accounts, AWS Region, network paths, identities, owners, workload boundaries, and isolation requirements. If you need guidance creating a cluster, refer to the [Amazon SageMaker HyperPod documentation](https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-hyperpod-quickstart.html) or the [AI on SageMaker HyperPod site](https://awslabs.github.io/ai-on-sagemaker-hyperpod).\n\nSeparate administrative and workload identities. Preserve the SageMaker HyperPod distinction between cluster administrators and data scientist users when creating project and access roles. A project-facing role shouldn’t receive cluster lifecycle permissions solely because its users run workloads. Refer to [AWS Identity and Access Management for SageMaker HyperPod](https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-hyperpod-prerequisites-iam.html) for the supported permission model.\n\nTreat networking as an end-to-end control. Restrict Amazon EKS Kubernetes API endpoint access to approved administrative network paths, control pod ingress and egress, and verify that the project, connection role, workload role, and cluster network support only the intended data paths. For orchestrator-specific configuration, refer to [Orchestrating SageMaker HyperPod clusters with Amazon EKS](https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-hyperpod-eks.html).\n\nFor example, configure private access to the Amazon EKS API endpoint and allow it only from approved administrator or workload subnets. Apply a default-deny Kubernetes `NetworkPolicy` and explicitly allow required service-to-service and egress paths. Use virtual private cloud (VPC) endpoints for services such as Amazon S3, Amazon Elastic Container Registry (Amazon ECR), and Amazon CloudWatch where appropriate, and give each tenant a dedicated workload role for data access.\n\nCreate a connection contract. A connection contract is a customer-managed governance record, such as a wiki page, ticket, service-catalog record, or file tracked in an infrastructure-as-code repository. It isn’t a platform feature. Record the following information for every approved project-to-cluster connection:\n\n- Business owner, operations owner, and cost owner.\n- SageMaker Unified Studio domain unit and project.\n- Cluster account, Region, name, and orchestrator.\n- Project role and the access role Amazon Resource Name (ARN) used by the connection.\n- Approved workload types and data classification.\n- EKS namespaces or Slurm access scope.\n- Scheduling policy and exception owner.\n- Monitoring, support, and decommissioning expectations.\n\nThis contract gives reviewers one record for evaluating and approving the complete access path.\n\n## Govern identity and task visibility\n\nControl both actions and visibility. Improper setup of task names, namespaces, resource requests, and usage patterns can reveal information about another team’s work.\n\nUse groups rather than individual grants for project membership and cluster access where possible. Review each role at its own boundary instead of creating one broad role that spans the domain, project, cluster, and workload policy.\n\nThe following workflow shows how to separate what a user can see from what a user can do.\n\nReview default task visibility before onboarding users. The [documentation for SageMaker HyperPod in SageMaker AI Studio](https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-hyperpod-studio-tabs.html) explains that SageMaker AI Studio users can see all Amazon EKS cluster tasks by default. For Slurm clusters, every SageMaker AI Studio user can view, manage, and interact with available tasks. Configure task-view restrictions before onboarding multiple teams: see [Restrict task view in Studio for EKS clusters](https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-hyperpod-studio-setup-eks.html#sagemaker-hyperpod-studio-setup-eks-restrict-tasks-view) for Amazon EKS and [Restrict task view in Studio for Slurm clusters](https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-hyperpod-studio-setup-slurm.html#sagemaker-hyperpod-studio-setup-slurm-restrict-tasks-view) for Slurm. Project membership isn’t a cluster-level security boundary.\n\nFor Amazon EKS clusters, map each team to an approved namespace and RBAC permissions, and map each workload service account to a tenant-specific IAM role. For cross-account data access, associate the service account with an EKS Pod Identity role in the capacity (cluster) account and set a target IAM role on the association (`targetRoleArn`). EKS Pod Identity then performs the cross-account role assumption automatically, so application code doesn’t need to call `AssumeRole`. The target role lives in the consumer or data account. Keep read-only visibility separate from permissions to create, update, or delete workloads. For Slurm clusters, define equivalent user, account, partition, file-system, and task-visibility controls.\n\nFor example, the following Kubernetes RBAC Role grants a team read-only access to jobs in only its own namespace. A separate role is required to create or delete workloads, which keeps “see” separate from “act.”\n\n## Translate business priorities into scheduling policy\n\nFor Amazon EKS clusters, apply task governance only after identity and workload access controls are in place.\n\nThe following figure separates the two decisions this involves.\n\nDocument guaranteed and shared capacity, priority classes, and whether a team can use another team’s idle allocation. Assign an owner to every exception. Task governance also applies to [SageMaker HyperPod spaces](https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-hyperpod-studio-tabs.html) (self-contained JupyterLab or Code Editor environments that run directly on the cluster), so include these interactive development workloads in the allocation policy.\n\nFor example, on an Amazon EKS cluster, you might give the production training team a guaranteed 60 percent allocation with a high priority class, let the research team borrow idle capacity at a lower priority, and allow preemption of that borrowed capacity when the production team submits work. Assign one owner to approve any exception to this policy so that one-time requests don’t quietly become the norm.\n\nKeep authorization and scheduling policy separate. EKS RBAC or Slurm ACLs determine whether a user can submit a workload. SageMaker HyperPod task governance for Amazon EKS, or native Slurm scheduling controls, determine when that authorized workload receives compute. Using scheduler quota as an access control, or using authorization controls as scheduling policy, produces unclear behavior and makes incidents harder to diagnose.\n\nThe post [Best practices for Amazon SageMaker HyperPod task governance](https://aws.amazon.com/blogs/machine-learning/best-practices-for-amazon-sagemaker-hyperpod-task-governance/) explains fair-share weights, quotas, lending and borrowing, priority classes, and common allocation scenarios. For Amazon EKS clusters, use those patterns when defining allocation policy. For Slurm clusters, use the native scheduler controls described earlier.\n\n## Use observability as a governance feedback loop\n\nWith SageMaker Unified Studio, you can view SageMaker HyperPod cluster details for tasks, metrics, settings, and metadata. For Amazon EKS clusters, task governance metrics include hardware, team, and task views. These views help you compare policy intent with actual consumption.\n\nTreat observability as a loop, as shown in the following figure.\n\nDefine an owner and a response for every signal you monitor. The following examples turn dashboard data into administrative decisions.\n\n| **Signal** | **Administrative decision** | \n| Cluster capacity and accelerator utilization | Determine whether low utilization is temporary, policy-driven, or caused by workload constraints | \n| Team allocation and utilization | Review whether reserved and shared capacity still reflects business demand | \n| Task run time and wait time | Investigate priority policy, workload sizing, or capacity contention | \n| Pending and preempted tasks | Confirm that scheduler outcomes match the approved priority model | \n| Node health and recovery events | Invoke the cluster incident process and validate recovery objectives | \n\nFor Amazon EKS clusters, the task table shows Kubeflow tasks (PyTorch, MPI, and TensorFlow), and PyTorch tasks are shown by default. Workloads submitted through other mechanisms might not appear there. For Slurm clusters, task tables show jobs in the current scheduler queue, while Slurm accounting provides historical job data through tools such as `sacct`. Define where you obtain historical task data, audit evidence, and incident details.\n\nThe Amazon CloudWatch Observability EKS add-on is required for the documented metrics views. The add-on has its own prerequisites: version 2.4.0 or later, and the `CloudWatchAgentServerPolicy`  IAM policy attached to the Kubernetes worker node role. Kueue metrics, which supply the task governance views, might incur CloudWatch metrics charges after the free tier. Refer to the [SageMaker HyperPod dashboard documentation](https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-hyperpod-eks-operate-console-ui-governance-metrics.html) for current setup and pricing considerations.\n\n## Review connections throughout their lifecycle\n\nSet a review schedule for every connection. Also define event-driven reviews for changes to project ownership, access roles, account or Region, cluster capacity, orchestrator version, data classification, or monitoring coverage. Use the connection contract to record each decision.\n\nThe following figure shows the lifecycle stages.\n\nAutomate inventory and evidence collection where it reduces manual work but keep approval with accountable administrators. Revoke connections that no longer have a business purpose or responsible owner.\n\n## Conclusion\n\nWith Amazon SageMaker Unified Studio, machine learning teams can follow a project-centered path to approved SageMaker HyperPod compute while infrastructure teams retain control of the centralized cluster. Start with one non-production cluster and a test project. Configure workload identity, namespace or Slurm scope, data permissions, network policy, task-view restrictions, and task governance for Amazon EKS or native scheduling controls for Slurm before adding members. Install the Amazon CloudWatch Observability EKS add-on so metrics views populate, and record the complete approval in the connection contract. Use cross-account roles for approved consumer accounts, and use dedicated infrastructure when hard isolation is required. Follow the [SageMaker HyperPod connection procedure](https://docs.aws.amazon.com/sagemaker-unified-studio/latest/userguide/sagemaker-hyperpods.html) for implementation steps.", "url": "https://wpnews.pro/news/best-practices-for-amazon-sagemaker-hyperpod-administration-and-governance", "canonical_source": "https://aws.amazon.com/blogs/machine-learning/best-practices-for-amazon-sagemaker-hyperpod-administration-and-governance/", "published_at": "2026-10-06 15:50:23+00:00", "updated_at": "2026-10-06 16:16:12.480677+00:00", "lang": "en", "topics": ["machine-learning", "mlops", "ai-infrastructure", "ai-policy"], "entities": ["Amazon", "Amazon SageMaker HyperPod", "Amazon SageMaker Unified Studio", "Amazon SageMaker AI", "Amazon Elastic Kubernetes Service", "AWS Identity and Access Management", "Slurm", "JupyterLab"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/best-practices-for-amazon-sagemaker-hyperpod-administration-and-governance", "markdown": "https://wpnews.pro/news/best-practices-for-amazon-sagemaker-hyperpod-administration-and-governance.md", "text": "https://wpnews.pro/news/best-practices-for-amazon-sagemaker-hyperpod-administration-and-governance.txt", "jsonld": "https://wpnews.pro/news/best-practices-for-amazon-sagemaker-hyperpod-administration-and-governance.jsonld"}}