Artificial Intelligence #
Multiple teams within the same company increasingly need shared access to expensive GPU clusters for their generative AI operations, while maintaining isolation boundaries, resource fairness, and operational independence. Consider a data science team training large language models, a computer vision group running inference workloads, and a research team experimenting with new model architectures. All of them might need access to the same cluster. Without a well-designed multi-tenant (multi-team) architecture, organizations face uncontrolled resource consumption, weak isolation between teams, an inability to attribute shared GPU costs to the teams that incur them, and administrative overhead that slows down innovation.
Amazon SageMaker HyperPod is a purpose-built AI service that simplifies the management of large-scale compute clusters for gen AI workloads. It provides resilient, optimized clusters orchestrated by Amazon Elastic Kubernetes Service (Amazon EKS) or Slurm, so organizations can run distributed training, interactive development, and model inference at scale. At the same time, it automatically handles node health monitoring, fault recovery, and cluster lifecycle management.
In this post, we present a reference architecture for building a multi-tenant environment on Amazon SageMaker HyperPod with EKS. This architecture uses AWS IAM Identity Center for centralized authentication, per-team SageMaker AI domains for a tailored user experience, Kubernetes namespaces for workload isolation, HyperPod Task Governance for fair resource allocation, and namespace-level cost allocation for per-team spend visibility and chargeback. By the end of this post, you will have a clear blueprint for multiple teams to efficiently share a single HyperPod EKS cluster.
Architecture overview #
The following diagram illustrates the high-level architecture of a multi-tenant HyperPod EKS deployment. In this example, two teams (Team A and Team B) share a single HyperPod EKS cluster, each operating within their own isolated namespace.
The architecture is structured as a layered flow from left to right, connecting user identity through authorization controls and into isolated workload namespaces on the cluster.
Users and authentication
On the far left, individual users from each team (User 1 from Team A, User 2 from Team B) interact with the system through two paths. Both paths authenticate through the AWS IAM Identity Center Portal, which federates with an external identity provider (such as Microsoft Entra ID) shown at the bottom-left.
The first path is through CLI access. Users authenticate with aws sso login, which redirects them to the Identity Center portal, and then obtain temporary credentials from their team’s permission set to submit tasks directly to the EKS cluster with kubectl. In the diagram, the pink arrow flows from the CLI terminal across the top directly into the HyperPod EKS cluster.
The second path is through the Identity Center portal directly, where users select the SageMaker Studio application to sign in to their team-specific SageMaker AI domain.
Each team has a corresponding permission set (TeamA permission set, TeamB permission set) that carries the AWS Identity and Access Management (IAM) policies needed for CLI workflows. Identity Center automatically provisions an IAM role for each permission set, shown in the diagram as TeamA-permissionset-role and TeamB-permissionset-role (labeled “CLI/Console role”). This role serves as the IAM principal when users authenticate through aws sso login.
SageMaker AI domains
From the Identity Center portal, users are routed to their team-specific SageMaker AI domain. Each domain (SageMaker AI domain Team A and SageMaker AI domain Team B) provides a dedicated Amazon SageMaker Studio GUI and is configured with a team-specific execution role (TeamA-role and TeamB-role respectively). These domains serve as the primary workspace interface, so users can submit tasks from the GUI (shown by the pink arrows flowing toward EKS).
EKS access control
At the EKS boundary, access entries map IAM roles to Kubernetes permissions. The diagram shows access entries for TeamA-role and TeamB-role (the Studio execution roles), which authorize requests originating from the SageMaker Studio GUI. Access entries must also be configured for the Identity Center-provisioned CLI/Console roles (TeamA-permissionset-role and TeamB-permissionset-role) to authorize requests arriving through kubectl. All access entries are associated with managed or custom role-based access control (RBAC) policies (represented by the key icon) and scoped to the team’s designated namespace. As a result, whether access originates from Studio or the CLI, users can only interact with resources in their own namespace.
HyperPod EKS cluster
The cluster itself is depicted with two cross-cutting platform layers at the top: HyperPod Observability (for monitoring and dashboards) and HyperPod Task Governance (for compute quota management and scheduling priorities). Below these layers, the cluster is partitioned into Namespace A (Team A) and Namespace B (Team B). Within each namespace, teams can run their own HyperPod Spaces (interactive development environments), HyperPod PyTorch jobs (distributed training workloads), and HyperPod Inference endpoints (model serving).
Storage
Beneath the cluster, the architecture includes two storage tiers. The first is a POSIX-compliant file system (Amazon FSx for Lustre or Amazon FSx for OpenZFS) organized into per-team shared directories (/fsx/TeamA, /fsx/TeamB) and per-user home directories (/home/User1, /home/User2). The second is per-team or shared Amazon Simple Storage Service (Amazon S3) buckets for object storage, governed by the team’s IAM execution role.
This architecture isolates each team from authentication through authorization to workload execution, while sharing expensive GPU infrastructure efficiently.
Authentication and access control #
The foundation of any multi-tenant system is robust authentication: verifying who users are before they interact with any resource. In this architecture, AWS IAM Identity Center serves as the centralized authentication layer, federating with an external identity provider to manage user identities and group memberships.
Why AWS IAM Identity Center
AWS IAM Identity Center (successor to AWS Single Sign-On) provides a single place to manage workforce identities across AWS accounts and applications. For a multi-tenant HyperPod deployment, it offers several key capabilities:
- Centralized identity management – Rather than maintaining separate user databases per AWS service, Identity Center provides a single source of truth for all user identities and their group memberships.
- Federation with existing identity providers – Most enterprises already manage their workforce identities in systems like Microsoft Entra ID (formerly Azure AD), Okta, or Ping Identity. Identity Center integrates with these providers, so organizations can reuse their existing identity infrastructure without duplicating user accounts.
- Native integration with SageMaker AI – SageMaker AI domains support Identity Center authentication, so users can sign in to SageMaker Studio through their corporate identity provider with single sign-on (SSO).
- AWS account access – Identity Center can also grant users access to the underlying AWS account with specific permission sets, which support CLI workflows alongside the Studio GUI experience.
- Required for Amazon Managed Grafana – Amazon Managed Grafana uses Identity Center as its authentication mechanism for workforce users, making it the natural choice when teams also need access to observability dashboards for monitoring their workloads.
Learn more: What is IAM Identity Center
Configuring Identity Center with an external identity provider
In this reference architecture, we use Microsoft Entra ID as the external identity provider, though the same pattern applies to most standard Security Assertion Markup Language (SAML) 2.0 providers.
The configuration involves:
- Group structure in the identity provider – In Entra ID, create groups that correspond to your organizational teams. In our example, we define three groups:
TeamA,TeamB, andAdmin. Each group contains the users belonging to that team (for example,user1-teamA@example.comin theTeamAgroup). - SCIM provisioning – Enable SCIM (System for Cross-domain Identity Management) synchronization between Entra ID and AWS IAM Identity Center. SCIM provides automatic provisioning and de-provisioning of users and groups. When a new user is added to the
TeamAgroup in Entra ID, they’re automatically synchronized to Identity Center and gain appropriate access without manual intervention. - SAML-based authentication – Configure SAML 2.0 federation so that when users authenticate, they do so against Entra ID. Identity Center acts as the service provider, trusting the assertions from your Entra ID tenant.
With this configuration, you manage team membership (which drives all downstream authorization decisions) in your existing corporate directory, and it propagates to AWS automatically.
The following image shows an example of how organizational teams can be represented in Microsoft Entra ID, with dedicated groups for TeamA, TeamB, and Admin.
Then, the following image shows the corresponding groups in AWS IAM Identity Center, automatically provisioned from Entra ID through SCIM synchronization.
Learn more: Connect an external identity provider · SCIM profile and SAML 2.0 implementation
Authorization #
With authentication established, the next layer is authorization: controlling what actions each team can perform across AWS services and the Kubernetes cluster. Authorization in this architecture operates at two levels: IAM for service-level access, and Kubernetes RBAC for cluster-level access.
Per-team IAM roles
Each team requires a dedicated IAM role that encapsulates the AWS level permissions needed for their AI and machine learning (ML) workflows. These roles serve as the SageMaker AI domain execution role and define what AWS services the team can access.
A typical team IAM role should include policies granting access to:
- Amazon SageMaker AI – For managing HyperPod clusters, MLflow tracking servers, and other SageMaker AI resources through the SageMaker AI API.
- Amazon S3 – For reading training datasets and writing model artifacts, checkpoints, and logs. Scope these permissions to team-specific bucket prefixes.
- Amazon CloudWatch – For viewing logs and metrics related to the team’s workloads.
- Amazon EKS – Specifically, the
eks:AccessKubernetesApiandeks:MutateViaKubernetesApipermissions, which the SageMaker Studio GUI needs to make Kubernetes API calls on behalf of the user (for example, listing Spaces or submitting jobs).
The trust policy on each IAM role must include sagemaker.amazonaws.com as a trusted principal, so SageMaker AI can assume the role on behalf of users when they operate through Studio. If you plan to reuse the same execution role as an EKS Pod Identity association for in-cluster workloads (as discussed later in the Amazon S3 storage section), the trust policy must also include pods.eks.amazonaws.com as a trusted principal. CLI access through Identity Center uses a separate permission set with its own policies (see the “AWS account access through Identity Center” section), so CLI permissions can be independently scoped.
Learn more: How to use SageMaker AI execution roles
AWS account access through Identity Center
Beyond SageMaker Studio, teams often need direct AWS account access for CLI operations such as running kubectl commands, scripting workflows, or accessing resources programmatically. Identity Center permission sets provide this capability.
For the **Admin** group, assign a permission set with administrative access as required by your company’s policies, granting the necessary account access for cluster management and administrative operations.
For **Team A** and **Team B**, create permission sets with inline or managed policies that grant the permissions needed for CLI workflows directly. A typical team permission set includes permissions for `eks:AccessKubernetesApi` (to view Kubernetes resources from AWS Console), scoped S3 access for the team’s data, and CloudWatch read access for monitoring. These policies are defined independently from the Studio execution role, so administrators can tailor CLI permissions to the specific operations that teams perform from the command line.
Users retrieve temporary credentials through the AWS Command Line Interface (AWS CLI) using aws sso login, which they can then use to configure kubectl for direct interaction with the EKS cluster.
The following image shows the per-team permission sets in AWS IAM Identity Center, providing scoped AWS account access for CLI workflows such as running kubectl and aws sso login against the EKS cluster.
Learn more: Manage AWS accounts with permission sets
Configuring the AWS CLI
Team members configure the AWS CLI to authenticate through Identity Center by running aws configure sso. This creates profiles in ~/.aws/config that reference the appropriate Identity Center session and permission set. Each team member uses their team-specific profile when interacting with the cluster from the command line, maintaining authorization boundaries whether access originates from Studio or from a local terminal.
The resulting configuration defines a shared sso-session block for the Identity Center portal and one named profile per team, each pointing at that team’s permission set. Team members then run aws sso login --profile <team> to obtain temporary credentials scoped to their permission set:
Learn more: Configuring IAM Identity Center authentication with the AWS CLI
SageMaker AI domains #
SageMaker AI domains provide the workspace boundary for each team, offering a tailored user experience, pre-configured execution roles, and built-in integration with Identity Center authentication.
Why SageMaker AI domains
Using one SageMaker AI domain per team is a well-established pattern for organizing multi-team environments. This approach offers several advantages:
- Established multi-team pattern – AWS has documented this approach extensively for separating lines of business or teams with multiple domains, making it a proven and supported configuration.
- Native Identity Center authentication – Each domain can be configured with Identity Center authentication, meaning users sign in once through their corporate identity provider and land directly in their team’s Studio environment.
- Built-in team configuration – Domains already provide mechanisms to specify configurations for users and teams without requiring additional custom entities. For example, settings like team execution role can be specified at domain level, and overridden at user profile level for maximum flexibility.
- Navigation customization – With domain settings, administrators can hide navigation items that are not relevant to the team’s workflows, presenting a focused interface tailored to HyperPod use cases.
Learn more: SageMaker AI domain entities and statuses · Multiple domains overview
Setting up per-team domains
Create one SageMaker AI domain per team with Identity Center authentication. In our example, we create TeamA-domain and TeamB-domain. Each domain is configured as follows:
- Default execution role – Set the domain’s default execution role to the team-specific IAM role created in the authorization step. All actions performed through Studio inherit the appropriate permissions as a result.
- Identity Center group assignment – Add the corresponding Identity Center group (for example, the
TeamAgroup) to the domain. This activates the SageMaker Studio application for all members of that group, granting them access to the Studio interface. - Application assignment verification – After configuring group access, review the application assignment in Identity Center to confirm that the correct groups are mapped to the correct domains.
- Navigation customization – Configure the default navigation settings for each domain to present only the relevant capabilities. For example, you might hide items not related to HyperPod workflows, providing a streamlinedHyperPod-focused user experience that reduces cognitive overhead for team members who only need to work with HyperPod resources.
The following image shows the SageMaker AI console with one domain per team (TeamA-domain and TeamB-domain), each providing an isolated workspace boundary.
Then the following image shows the details of TeamA-domain, including the assigned Identity Center groups.
HyperPod EKS cluster configuration #
The HyperPod EKS cluster is where workloads are executed. Multi-tenancy at the cluster level is achieved through Kubernetes namespaces for isolation and EKS access entries for authorization.
Namespace isolation
Create a dedicated Kubernetes namespace for each team, for example hyperpod-ns-team-a and hyperpod-ns-team-b. Namespaces provide a logical boundary within the cluster, isolating each team’s workloads (Spaces, training jobs, inference endpoints) from one another.
Note: Namespaces are an isolation boundary, not a hard security boundary. This architecture targets multi-team within a single organization: teams that share a cluster under a common administrative domain and a baseline of mutual trust. It isn’t designed for multi-customer isolation between mutually untrusted tenants.
Namespaces, RBAC, and quotas prevent accidental interference (teams overwriting each other’s resources or exceeding their compute allocation) but aren’t a defense against a determined malicious tenant: namespaced pods share the same nodes and kernel, and cluster-scoped resources (nodes, PersistentVolumes, CRDs, some operator components) sit outside any namespace.
For untrusted tenants or strict regulatory isolation, use stronger boundaries such as separate clusters or accounts, dedicated node pools, and runtime sandboxing. For the multi-team scenario here, namespace isolation combined with RBAC, Task Governance quotas, and the POSIX identity controls described later strike an appropriate balance of separation and operational simplicity.
Namespaces can be created manually with kubectl create namespace or provisioned automatically through HyperPod Task Governance, which manages namespaces as part of its quota and scheduling configuration.
The following image shows the cluster namespaces (managed with HyperPod Task Governance), with one dedicated namespace per team (hyperpod-ns-team-a and hyperpod-ns-team-b) providing workload isolation.
Learn more: Kubernetes namespaces
Network isolation
Namespaces don’t restrict network traffic. By default, Kubernetes networking is flat: every pod can reach every other pod across all namespaces. As a result, a pod in hyperpod-ns-team-a can open a connection to a pod in hyperpod-ns-team-b unless you add controls. To scope pod-to-pod reachability along team boundaries, use Kubernetes NetworkPolicy resources.
The recommended pattern is default-deny per namespace: start by denying all ingress (and optionally egress), then explicitly allow the traffic each team needs, typically intra-namespace communication plus required egress such as DNS, storage endpoints, and AWS APIs. The following example denies all ingress in a team’s namespace and then allows traffic only from pods within the same namespace:
NetworkPolicy enforcement depends on a Container Network Interface (CNI) that supports it. On EKS, you can enable network policy support in the Amazon Virtual Private Cloud (Amazon VPC) CNI.
As with namespaces, NetworkPolicies reduce accidental cross-team reachability and shrink the scope, but they are not by themselves an adversarial security boundary on shared nodes. For stronger separation, consider dedicated node pools per team or separate clusters, as noted earlier.
Learn more: Kubernetes network policies · Amazon VPC CNI network policy
EKS access entries
EKS access entries connect IAM principals to Kubernetes RBAC permissions. For each team, create two access entries:
- Studio access entry – The IAM principal is the team’s SageMaker AI domain execution role. This entry is used when actions originate from the SageMaker Studio GUI.
- CLI access entry – The IAM principal is the SSO-provisioned role created by Identity Center for the team’s permission set (following the pattern
AWSReservedSSO_<permission-set-name>_<unique-id>). This entry is used when users interact with the cluster throughkubectl.
Both entries are scoped to the team’s namespace with managed or custom Kubernetes policies. For example, both Team A entries grant permissions only within hyperpod-ns-team-a. The two entries can carry different RBAC policies if you want. For instance, the CLI entry might restrict write access to certain resource types while the Studio entry allows full access.
With this scoping, whether access originates from Studio or from the CLI, users can only interact with resources in their own namespace. Attempting to list or modify resources in another team’s namespace results in a Kubernetes Forbidden error.
For more advanced scenarios, you can use Kubernetes groups in the access entry to map users to custom ClusterRoles or Roles that provide fine-grained permissions beyond the standard managed policies. The following image shows an EKS access entry for Team B’s role, scoped down to the hyperpod-ns-team-b namespace, so its permissions apply only within Team B’s namespace.
Learn more: Grant IAM users access to Kubernetes with EKS access entries
HyperPod Task Governance
When Task Governance is enabled on the cluster, it provides an additional layer of resource management:
- Compute quotas – Define how much GPU and CPU capacity each team can consume. This prevents a single team from monopolizing shared hardware during training runs.
- Priorities – Assign scheduling priorities to each team or workload type, allowing critical production inference workloads to preempt experimental training jobs when resources are constrained.
- Fair scheduling – With Task Governance, when multiple teams are competing for resources, allocation follows the configured policies rather than a first-come-first-served model.
Configure Task Governance with appropriate quotas and priorities per team namespace, balancing between guaranteed minimum allocations and burst capacity for bursty workloads.
The following image shows the Task Governance compute allocations for the two teams, with each team’s namespace assigned its own quota of cluster compute capacity.
Learn more: SageMaker HyperPod task governance
Storage #
Storage is a key component of artificial intelligence and machine learning (AI/ML) environments shared across teams. Teams need high-performance file systems for training data, checkpoints, and model artifacts, while maintaining appropriate access boundaries between teams.
POSIX-compliant file systems
For workloads that require a shared, high-performance POSIX file system (common for distributed training where multiple nodes read the same dataset or write checkpoints), consider the following options:
- Amazon FSx for Lustre – Provides high-throughput, low-latency parallel file system access, ideal for large-scale training workloads that need to read large datasets at high speed.
- Amazon FSx for OpenZFS – Offers a general-purpose file system with strong POSIX semantics, snapshots, and compression. Well-suited for workloads that need traditional file system features alongside high performance.
- Amazon Elastic File System (Amazon EFS) – Provides fully managed, elastic Network File System (NFS) storage. EFS also supports access points, which can simplify per-team directory isolation by mapping different mount points to different directories with enforced UIDs and GIDs.
The storage layout typically follows this structure:
- Per-team shared directories – Each team has a shared directory (for example,
/fsx/TeamA,/fsx/TeamB) for datasets, models, and artifacts that all team members need to access. - Per-user home directories – Each user has a personal home directory (for example,
/home/User1,/home/User2) for individual work, experiments, and notebooks.
The POSIX permission model on these file systems relies on UIDs, GIDs, and supplemental groups to enforce access boundaries. These POSIX identities should then be propagated to the pod security context when a user launches a HyperPod Space or submits a training job, so that file system access respects the configured ownership and permissions. We recommend using a Kubernetes mutating admission webhook to retrieve POSIX identity information from your identity store. When a workload is submitted, the webhook looks up the identity at runtime and modifies the security context of the Pod accordingly.
Learn more: FSx for Lustre · FSx for OpenZFS · Amazon EFS
Amazon S3 storage
For object storage, access to S3 buckets is governed by the team’s IAM execution role. You can create per-team buckets or use a shared bucket with per-team prefixes, relying on IAM policies to enforce isolation. Pods within the cluster need appropriate service accounts configured with IAM Roles for Service Accounts (IRSA) or Pod Identity to authenticate to S3. For simplicity, you can associate the same execution role configured on the SageMaker AI domain to a Kubernetes service account within the team’s namespace, providing consistent S3 access from both Studio and cluster workloads. Learn more: IAM roles for service accounts (IRSA) · EKS Pod Identity
HyperPod Spaces #
HyperPod Spaces provide interactive development environments (IDEs) running directly on cluster nodes. On a shared cluster, Spaces must be properly scoped to each team’s namespace and configured with appropriate resource templates.
Space templates
Create namespace-scoped Space templates for each team. These templates define the resource configurations (instance types, storage volumes, environment variables) available to team members when creating Spaces. By scoping templates to a namespace, you make sure that each team can only launch Spaces within their designated boundary.
When HyperPod Task Governance is enabled, templates should include the default labels required by the governance system (such as team identifiers and priority labels). The cluster administrator pre-configures these labels so that team members don’t need to specify them manually when launching Spaces.
The following example shows a JupyterLab Space template scoped to Team A. The team-specific parts are the metadata.namespace, the Task Governance queue label under baseLabels, and the defaultVolumes that mount the team’s shared file system and the user’s home directory:
Persistent volume claims
Create the appropriate Persistent Volume Claims (PVCs) in each team’s namespace, referencing the shared file system. These PVCs mount the team’s shared directory and the user’s home directory into the Space, providing access to training data, checkpoints, and personal workspaces.
Owner-only and shared Spaces
Consider your organization’s requirements around Space sharing:
- Owner-only Spaces – Each Space is accessible only by the user who created it. This is the default configuration, appropriate when teams work on sensitive or independent projects.
- Shared Spaces – Multiple team members can access the same Space, useful for pair programming, collaborative debugging, or shared development environments. When enabling shared Spaces, make sure that POSIX permissions and supplemental groups are configured to allow appropriate access to files created within the Space.
Learn more: Interactive development environments on Amazon SageMaker HyperPod EKS clusters
Studio experience #
While teams can interact with the cluster entirely from the CLI using kubectl, SageMaker Studio provides a graphical entry point to the cluster for users who prefer a managed, GUI-driven workflow. In this architecture, each team accesses Studio through its own SageMaker AI domain (as described earlier), signing in with the same Identity Center credentials and operating within the boundaries of the team’s namespace.
The following image shows the IAM Identity Center access portal that users reach after signing in with their corporate credentials, providing single sign-on access to their assigned SageMaker Studio applications and Amazon Managed Grafana applications.
From the Studio UI, team members can:
- Manage HyperPod Spaces – Launch interactive development environments from the namespace-scoped Space templates the administrator has configured, without writing Kubernetes manifests or specifying Task Governance labels manually. Team members can also start, stop, and connect to their running Spaces, opening the associated IDE (such as JupyterLab) directly in the browser.
- Manage Ray workloads – Create and monitor Ray clusters, connect a JupyterLab or Code Editor workspace to a cluster, submit distributed jobs, and open the Ray Dashboard and Amazon Managed Grafana observability dashboards, all without writing Kubernetes manifests or running
kubectlcommands.
Because Studio operates through the team’s Domain execution role and the corresponding EKS access entry, all actions are scoped to the team’s namespace. A user launching a Space or a Ray cluster from Studio can only create it within their own team’s boundary, consistent with the isolation model enforced for CLI access.
The following image shows how to create a HyperPod Space from the SageMaker Studio UI, where a team member selects a namespace-scoped Space template without writing Kubernetes manifests or specifying Task Governance labels manually.
Learn more: Interactive development environments on Amazon SageMaker HyperPod EKS clusters · Introducing new Ray capabilities on SageMaker HyperPod
HyperPod Training Operator #
The HyperPod Training Operator enables teams to submit distributed training jobs as Kubernetes custom resources (for example, HyperPodPyTorchJob). In the multi-tenant architecture, training jobs are namespace-scoped, meaning they automatically inherit the team’s isolation boundaries.
Teams can submit training jobs from the CLI using kubectl apply with the appropriate job manifest. The job runs in the team’s namespace, uses the team’s compute quotas (if Task Governance is enabled), and has access to the team’s storage volumes.
When Task Governance is enabled, training jobs are subject to the team’s allocated quotas and priority settings. If a team has consumed its guaranteed allocation, jobs might be queued until resources become available or until lower-priority workloads are preempted.
The two elements that anchor a job to a team are the metadata.namespace (which scopes the job to the team’s isolation boundary) and the Task Governance labels. Task Governance is built on Kueue, so the job is routed to the team’s local queue through kueue.x-k8s.io/queue-name. It’s assigned a scheduling priority through kueue.x-k8s.io/priority-class, whose value is the name of a WorkloadPriorityClass defined on the cluster:
Learn more: Using the HyperPod training operator
HyperPod Inference Operator #
The HyperPod Inference Operator allows teams to deploy models as inference endpoints directly on the cluster. Similar to training jobs, inference endpoints are namespace-scoped and subject to the team’s RBAC policies and Task Governance quotas.
Teams can deploy models from the CLI by creating inference endpoint custom resources in their namespace. The endpoints are isolated per namespace, meaning Team A can’t access or interfere with Team B’s inference endpoints.
For production inference workloads that require high availability, consider assigning higher scheduling priority to inference endpoints than to training jobs, so that model serving isn’t interrupted by batch training workloads.
As with training jobs, the inference endpoint is placed in the team’s metadata.namespace and carries the Task Governance labels. Here, the kueue.x-k8s.io/priority-class references a higher-priority WorkloadPriorityClass so that model serving can preempt batch training when the team’s resources are constrained:
Learn more: Deploying models on Amazon SageMaker HyperPod
HyperPod Observability #
Visibility into cluster health, workload performance, and resource utilization is essential for all teams. HyperPod Observability provides built-in monitoring and dashboarding capabilities through Amazon Managed Grafana.
Configuring team access to Grafana
Teams need access to observability dashboards to monitor their workloads, troubleshoot performance issues, and understand resource consumption. However, in a multi-tenant environment, this access should typically be read-only:
- Configure Identity Center authentication for Amazon Managed Grafana – On the Amazon Managed Grafana console, navigate to Authentication and enable AWS IAM Identity Center. Users can then sign in to Grafana with the same corporate credentials they use for SageMaker Studio.
- Assign team groups as Viewers – Map the Identity Center groups (
TeamA,TeamB) to the Grafana Viewer role. This grants team members read-only access to dashboards and metrics without the ability to modify dashboards or data sources. - Admin access – Assign the
Admingroup to the Grafana Admin or Editor role, so they can create and modify dashboards, configure alerting, and manage data sources. - Team-specific dashboards – Consider creating dedicated dashboards that filter data by namespace, so each team sees only their own workload metrics. Amazon Managed Grafana supports Grafana Teams (a Grafana-native RBAC concept, distinct from the organizational teams in this architecture), which can be mapped from Identity Center groups to restrict dashboard visibility and provide an additional layer of data isolation.
The following image shows the Grafana role assignments in Amazon Managed Grafana. The team groups (TeamA, TeamB) are assigned the Viewer role for read-only access to dashboards and metrics, while the admin group is assigned the Admin role, granting it the ability to create and modify dashboards, configure alerting, and manage data sources.
Learn more: Observability for Amazon SageMaker HyperPod cluster orchestrated by Amazon EKS
Cost allocation and chargeback #
In a multi-tenant environment where teams share expensive GPU infrastructure, understanding who consumes what is essential for accountability, budgeting, and chargeback. Kubecost addresses this need by breaking down in-cluster spend across native Kubernetes concepts (namespace, label, deployment, and service) and mapping it to organizational concepts like team, project, or environment.
Because this architecture already isolates each team in a dedicated namespace (hyperpod-ns-team-a, hyperpod-ns-team-b), namespace-level cost allocation aligns directly with team boundaries. This gives platform administrators a clear per-team view of GPU, CPU, memory, storage, and network consumption without any additional workload tagging. For step-by-step instructions on deploying and configuring Kubecost on a HyperPod cluster, see Kubecost on SageMaker HyperPod.
Enabling team visibility
After Kubecost is collecting data, group cost by namespace in the Allocations dashboard to see per-team spend. Because each team owns a namespace, this directly produces a per-team cost breakdown covering compute, memory, storage, and network. Just as with observability dashboards, teams benefit from visibility into their own cost data:
- Scope views to each team’s namespace – Kubecost supports filtering and saved reports by namespace, so each team can review its own consumption and trends without seeing other teams’ data.
- Set budgets and alerts – Configure per-namespace budget thresholds and alerts, so teams and platform administrators are notified when spending approaches defined limits, supporting the same resource-fairness goals as HyperPod Task Governance.
- Support chargeback and showback – Namespace-level allocation reports can feed internal chargeback (billing teams for their usage) or showback (reporting usage without billing) processes, giving finance and platform teams the data that they need to attribute shared GPU costs fairly.
The following image shows the Kubecost Allocations dashboard grouped by namespace, showing the cumulative cost over the last 7 days for each team’s namespace.
Learn more: Kubecost on SageMaker HyperPod · Kubecost
Conclusion #
This post presented a reference architecture for building multi-tenant environments on Amazon SageMaker HyperPod with EKS. By combining AWS IAM Identity Center for authentication, per-team IAM roles for AWS level authorization, SageMaker AI domains for a tailored workspace experience, Kubernetes namespaces for workload isolation, HyperPod Task Governance for fair resource allocation, and namespace-level cost allocation for per-team spend visibility, multiple teams can efficiently share a single HyperPod EKS cluster.
This is a flexible, composable approach that combines multiple building blocks into a cohesive solution. The architecture accommodates a variety of use cases and organizational structures. For example, organizations can extend this pattern to connect teams to HyperPod Slurm clusters alongside EKS, providing a unified multi-tenant experience across different orchestration backends.
While this approach requires assembling and configuring several components, the result is a high degree of control and customization that can be tailored to each organization’s specific isolation, compliance, and operational requirements. The foundational patterns (identity federation, namespace isolation, RBAC, quota-based governance, and cost allocation) will remain applicable. To get started, try building this multi-tenant setup on your own Amazon SageMaker HyperPod EKS cluster, and adapt the building blocks to your organization’s isolation, governance, and cost-allocation requirements.