Best practices for Amazon SageMaker HyperPod administration and governance Amazon published best practices for administering Amazon SageMaker HyperPod through Amazon SageMaker Unified Studio, laying out four layers of control — organization, project, cluster, and workload — for ML teams sharing accelerated compute. The guidance states that a SageMaker HyperPod connection adds an approved cluster to a project but does not replace the cluster's AWS Identity and Access Management (IAM), Amazon Elastic Kubernetes Service (Amazon EKS), or Slurm controls. The model is intended to let infrastructure teams keep cluster operations while offering approved compute to ML teams in their project context. Artificial Intelligence https://aws.amazon.com/blogs/machine-learning/ Best practices for Amazon SageMaker HyperPod administration and governance Amazon SageMaker HyperPod https://aws.amazon.com/sagemaker/hyperpod/ gives machine learning ML teams access to large pools of accelerated compute for training and fine-tuning models. When several teams share one cluster, the technical setup is usually straightforward. The challenging part is governance. You must decide which teams can use the cluster, how much capacity each team gets, what happens when one team’s workload competes with another’s, and who is accountable when usage drifts from policy. Amazon SageMaker Unified Studio https://aws.amazon.com/sagemaker/unified-studio/ adds another consideration: You can connect a SageMaker HyperPod cluster to a project so team members can launch workloads from their project workspace. That convenience is valuable, but after multiple teams share visibility into the same cluster, the controls that govern who can do what become even more important. In this post, we show how to administer SageMaker HyperPod through SageMaker Unified Studio while preserving the underlying governance controls. We cover the four layers of control: organization, project, cluster, and workload. We also explain how to design identity, capacity, and observability policies across them. By the end, you will have a repeatable model for offering approved SageMaker HyperPod compute to ML teams in their project context while keeping cluster operations with the infrastructure team. Amazon SageMaker HyperPod is a capability of Amazon SageMaker AI https://aws.amazon.com/sagemaker/ai/ . Amazon SageMaker Unified Studio is the data and AI development environment where teams build with their data and tools. With SageMaker Unified Studio, you can connect a project to an existing SageMaker HyperPod cluster. Members can then launch machine learning workloads, review cluster and task information, and open a JupyterLab workflow. You continue to manage clusters through Amazon SageMaker AI interfaces and APIs. This separation gives infrastructure teams a useful operating model. You can manage cluster infrastructure through established cloud operations processes. At the same time, you can present approved compute to machine learning teams in their project context. In this post, we describe how you can design infrastructure boundaries, govern access, allocate shared capacity, and operate SageMaker HyperPod consistently through SageMaker Unified Studio. Understand the administrative boundaries A well-governed environment separates organizational administration, project access, and cluster operations. Each layer answers a different question and uses a different control. The following table summarizes each boundary, its primary controls, and its administrative purpose. | Boundary | Primary controls | Administrative purpose | | Organization | SageMaker Unified Studio domains, domain units, associated accounts, project profiles, and authorization policies | Determines who can create projects, which accounts and Regions projects can use, and which tools are available | | Project | Project membership, project roles, and SageMaker HyperPod connections | Defines the collaboration context and the AWS resources that project members can access | | Cluster | SageMaker HyperPod cluster admin roles, Amazon EKS access entries, role-based access control RBAC , and EKS Pod Identity, or Slurm controls | Governs cluster configuration, scheduler access, namespaces, tasks, and infrastructure operations | | Workload | Compute allocations, priority classes, lending and borrowing policies, and task permissions | Controls who can submit work and how shared capacity is assigned | Treat these controls as layers. A SageMaker HyperPod connection adds an approved cluster to a project. It doesn’t replace the cluster’s AWS Identity and Access Management IAM , Amazon Elastic Kubernetes Service Amazon EKS , or Slurm controls. Review the project role, connection access role, EKS access entries and RBAC or Slurm controls, workload identity, data and AWS Key Management Service AWS KMS key https://aws.amazon.com/kms/ policies, network policy, task-view restrictions, and scheduler policy together. Make the cluster available only after those controls are aligned. These controls form four layers, shown in the following figure. SageMaker Unified Studio projects are collaboration boundaries. They aren’t strong runtime security boundaries. Keep the SageMaker HyperPod cluster, scheduler, and scarce accelerator capacity under one designated capacity account. Approved users and datasets can remain in the same account or in separate consumer and data accounts. AWS documents multi-account support for SageMaker HyperPod task governance on Amazon EKS clusters. Although On-Demand Capacity Reservations can be shared across accounts, centralize SageMaker HyperPod cluster ownership and capacity administration in the capacity account. Use approved cross-account access rather than making each consumer account an independent capacity administrator. - For Amazon EKS, use a namespace per tenant together with RBAC, per-tenant service accounts and EKS Pod Identity roles, default-deny network policies, and tenant-specific storage and AWS KMS permissions. For Slurm, use Slurm accounting with hierarchical accounts and associations, quality of service QoS , priority and fair-share policies, and partitions. Also use operating-system identity and file permissions, network controls, and tenant-specific data paths. - Use IAM roles and resource policies, rather than project membership alone, to control access to Amazon Simple Storage Service Amazon S3 buckets, AWS KMS keys, secrets, container registries, and other data services. Use project and domain-unit policies for collaboration and delegation. - For Amazon EKS clusters, use SageMaker HyperPod task governance, including quotas, priority classes, lending and borrowing, and preemption, to share the central accelerator pool fairly. For Slurm clusters, use native partitions, quality of service QoS , priority, fair-share, and preemption controls. Slurm doesn’t provide the same lending-and-borrowing model. These scheduling controls determine when an authorized workload receives compute. They don’t grant namespace or data access. For stronger isolation within a cluster, use dedicated nodes or node groups and admission controls where appropriate. Use a separate cluster or account when legal, regulatory, or security requirements demand hard infrastructure isolation. Cross-account consumer access can preserve account-level ownership without duplicating the central SageMaker HyperPod cluster. Within the capacity account, use project profiles https://docs.aws.amazon.com/sagemaker-unified-studio/latest/adminguide/custom.html and domain units https://docs.aws.amazon.com/sagemaker-unified-studio/latest/adminguide/domain-units.html with project authorization policies to delegate project creation and ownership. The following figure shows centralized capacity with tenant-specific workload, identity, data, network, and visibility controls. Figure 3 shows how those boundaries connect when SageMaker Unified Studio exposes approved SageMaker HyperPod compute to project members. Decide when SageMaker Unified Studio is the right administrative experience With SageMaker Unified Studio, machine learning teams that already work in a project can follow an approved path to shared SageMaker HyperPod compute. Members can find connected clusters, review status and metadata, inspect supported tasks and metrics, and move into JupyterLab without using a separate infrastructure inventory. This experience doesn’t replace cluster administration. The following table shows how each persona should use SageMaker Unified Studio alongside the service-specific interfaces. | Persona | Use SageMaker Unified Studio for | Continue using service-specific tools for | | Domain or infrastructure administrator | Domain units, project creation policy, project profiles, membership policy, and account placement | Organization-level controls, account provisioning, and infrastructure automation | | SageMaker HyperPod cluster administrator | Presenting approved cluster connections and reviewing cluster, task, settings, and metadata views | Cluster creation, updates, resiliency configuration, add-ons, EKS or Slurm administration, and incident response | | Project owner | Managing project membership and giving users a consistent project context for approved compute | Requesting infrastructure changes and approving business-specific access requirements | | ML engineer or data scientist | Finding approved compute, reviewing workload status, and opening the JupyterLab workflow | Submitting and managing detailed workloads through the SageMaker HyperPod CLI https://docs.aws.amazon.com/sagemaker/latest/dg/getting-started-hyperpod-training-deploying-models.html , kubectl , or Slurm tools as appropriate | Use SageMaker Unified Studio when project membership, data access, development tools, and compute need a common context. For cluster changes and repeatable automation, continue using SageMaker AI APIs, infrastructure as code, and orchestrator tools. For reference implementations and integrations, refer to the AI on SageMaker HyperPod https://awslabs.github.io/ai-on-sagemaker-hyperpod site. Make infrastructure decisions before connecting a cluster Before approving a connection, document the capacity account, any consumer or data accounts, AWS Region, network paths, identities, owners, workload boundaries, and isolation requirements. If you need guidance creating a cluster, refer to the Amazon SageMaker HyperPod documentation https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-hyperpod-quickstart.html or the AI on SageMaker HyperPod site https://awslabs.github.io/ai-on-sagemaker-hyperpod . Separate administrative and workload identities. Preserve the SageMaker HyperPod distinction between cluster administrators and data scientist users when creating project and access roles. A project-facing role shouldn’t receive cluster lifecycle permissions solely because its users run workloads. Refer to AWS Identity and Access Management for SageMaker HyperPod https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-hyperpod-prerequisites-iam.html for the supported permission model. Treat networking as an end-to-end control. Restrict Amazon EKS Kubernetes API endpoint access to approved administrative network paths, control pod ingress and egress, and verify that the project, connection role, workload role, and cluster network support only the intended data paths. For orchestrator-specific configuration, refer to Orchestrating SageMaker HyperPod clusters with Amazon EKS https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-hyperpod-eks.html . For example, configure private access to the Amazon EKS API endpoint and allow it only from approved administrator or workload subnets. Apply a default-deny Kubernetes NetworkPolicy and explicitly allow required service-to-service and egress paths. Use virtual private cloud VPC endpoints for services such as Amazon S3, Amazon Elastic Container Registry Amazon ECR , and Amazon CloudWatch where appropriate, and give each tenant a dedicated workload role for data access. Create a connection contract. A connection contract is a customer-managed governance record, such as a wiki page, ticket, service-catalog record, or file tracked in an infrastructure-as-code repository. It isn’t a platform feature. Record the following information for every approved project-to-cluster connection: - Business owner, operations owner, and cost owner. - SageMaker Unified Studio domain unit and project. - Cluster account, Region, name, and orchestrator. - Project role and the access role Amazon Resource Name ARN used by the connection. - Approved workload types and data classification. - EKS namespaces or Slurm access scope. - Scheduling policy and exception owner. - Monitoring, support, and decommissioning expectations. This contract gives reviewers one record for evaluating and approving the complete access path. Govern identity and task visibility Control both actions and visibility. Improper setup of task names, namespaces, resource requests, and usage patterns can reveal information about another team’s work. Use groups rather than individual grants for project membership and cluster access where possible. Review each role at its own boundary instead of creating one broad role that spans the domain, project, cluster, and workload policy. The following workflow shows how to separate what a user can see from what a user can do. Review default task visibility before onboarding users. The documentation for SageMaker HyperPod in SageMaker AI Studio https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-hyperpod-studio-tabs.html explains that SageMaker AI Studio users can see all Amazon EKS cluster tasks by default. For Slurm clusters, every SageMaker AI Studio user can view, manage, and interact with available tasks. Configure task-view restrictions before onboarding multiple teams: see Restrict task view in Studio for EKS clusters https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-hyperpod-studio-setup-eks.html sagemaker-hyperpod-studio-setup-eks-restrict-tasks-view for Amazon EKS and Restrict task view in Studio for Slurm clusters https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-hyperpod-studio-setup-slurm.html sagemaker-hyperpod-studio-setup-slurm-restrict-tasks-view for Slurm. Project membership isn’t a cluster-level security boundary. For Amazon EKS clusters, map each team to an approved namespace and RBAC permissions, and map each workload service account to a tenant-specific IAM role. For cross-account data access, associate the service account with an EKS Pod Identity role in the capacity cluster account and set a target IAM role on the association targetRoleArn . EKS Pod Identity then performs the cross-account role assumption automatically, so application code doesn’t need to call AssumeRole . The target role lives in the consumer or data account. Keep read-only visibility separate from permissions to create, update, or delete workloads. For Slurm clusters, define equivalent user, account, partition, file-system, and task-visibility controls. For example, the following Kubernetes RBAC Role grants a team read-only access to jobs in only its own namespace. A separate role is required to create or delete workloads, which keeps “see” separate from “act.” Translate business priorities into scheduling policy For Amazon EKS clusters, apply task governance only after identity and workload access controls are in place. The following figure separates the two decisions this involves. Document guaranteed and shared capacity, priority classes, and whether a team can use another team’s idle allocation. Assign an owner to every exception. Task governance also applies to SageMaker HyperPod spaces https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-hyperpod-studio-tabs.html self-contained JupyterLab or Code Editor environments that run directly on the cluster , so include these interactive development workloads in the allocation policy. For example, on an Amazon EKS cluster, you might give the production training team a guaranteed 60 percent allocation with a high priority class, let the research team borrow idle capacity at a lower priority, and allow preemption of that borrowed capacity when the production team submits work. Assign one owner to approve any exception to this policy so that one-time requests don’t quietly become the norm. Keep authorization and scheduling policy separate. EKS RBAC or Slurm ACLs determine whether a user can submit a workload. SageMaker HyperPod task governance for Amazon EKS, or native Slurm scheduling controls, determine when that authorized workload receives compute. Using scheduler quota as an access control, or using authorization controls as scheduling policy, produces unclear behavior and makes incidents harder to diagnose. The post Best practices for Amazon SageMaker HyperPod task governance https://aws.amazon.com/blogs/machine-learning/best-practices-for-amazon-sagemaker-hyperpod-task-governance/ explains fair-share weights, quotas, lending and borrowing, priority classes, and common allocation scenarios. For Amazon EKS clusters, use those patterns when defining allocation policy. For Slurm clusters, use the native scheduler controls described earlier. Use observability as a governance feedback loop With SageMaker Unified Studio, you can view SageMaker HyperPod cluster details for tasks, metrics, settings, and metadata. For Amazon EKS clusters, task governance metrics include hardware, team, and task views. These views help you compare policy intent with actual consumption. Treat observability as a loop, as shown in the following figure. Define an owner and a response for every signal you monitor. The following examples turn dashboard data into administrative decisions. | Signal | Administrative decision | | Cluster capacity and accelerator utilization | Determine whether low utilization is temporary, policy-driven, or caused by workload constraints | | Team allocation and utilization | Review whether reserved and shared capacity still reflects business demand | | Task run time and wait time | Investigate priority policy, workload sizing, or capacity contention | | Pending and preempted tasks | Confirm that scheduler outcomes match the approved priority model | | Node health and recovery events | Invoke the cluster incident process and validate recovery objectives | For Amazon EKS clusters, the task table shows Kubeflow tasks PyTorch, MPI, and TensorFlow , and PyTorch tasks are shown by default. Workloads submitted through other mechanisms might not appear there. For Slurm clusters, task tables show jobs in the current scheduler queue, while Slurm accounting provides historical job data through tools such as sacct . Define where you obtain historical task data, audit evidence, and incident details. The Amazon CloudWatch Observability EKS add-on is required for the documented metrics views. The add-on has its own prerequisites: version 2.4.0 or later, and the CloudWatchAgentServerPolicy IAM policy attached to the Kubernetes worker node role. Kueue metrics, which supply the task governance views, might incur CloudWatch metrics charges after the free tier. Refer to the SageMaker HyperPod dashboard documentation https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-hyperpod-eks-operate-console-ui-governance-metrics.html for current setup and pricing considerations. Review connections throughout their lifecycle Set a review schedule for every connection. Also define event-driven reviews for changes to project ownership, access roles, account or Region, cluster capacity, orchestrator version, data classification, or monitoring coverage. Use the connection contract to record each decision. The following figure shows the lifecycle stages. Automate inventory and evidence collection where it reduces manual work but keep approval with accountable administrators. Revoke connections that no longer have a business purpose or responsible owner. Conclusion With Amazon SageMaker Unified Studio, machine learning teams can follow a project-centered path to approved SageMaker HyperPod compute while infrastructure teams retain control of the centralized cluster. Start with one non-production cluster and a test project. Configure workload identity, namespace or Slurm scope, data permissions, network policy, task-view restrictions, and task governance for Amazon EKS or native scheduling controls for Slurm before adding members. Install the Amazon CloudWatch Observability EKS add-on so metrics views populate, and record the complete approval in the connection contract. Use cross-account roles for approved consumer accounts, and use dedicated infrastructure when hard isolation is required. Follow the SageMaker HyperPod connection procedure https://docs.aws.amazon.com/sagemaker-unified-studio/latest/userguide/sagemaker-hyperpods.html for implementation steps.