Learn learning medium confidence

How AWS Designs Shared GPU Clusters for Separate AI Teams

A HyperPod reference architecture combines centralized identity, Kubernetes namespaces, governed scheduling and per-team cost allocation on one EKS cluster.

Edited by Tyronne Panaino

AWS published a reference architecture on October 8 for organizations that want several AI teams to share one Amazon SageMaker HyperPod cluster without treating the cluster as one undifferentiated workspace. The design places Amazon EKS underneath a set of identity, namespace, scheduling and accounting controls.

The practical problem is broader than giving more users access to GPUs. A shared cluster has to distinguish who can submit work, which Kubernetes resources they can reach, how concurrent workloads receive capacity and how usage is attributed. AWS's blueprint treats those as separate control layers that must work together.

One cluster, separate team paths

The AWS architecture guide starts with AWS IAM Identity Center as the common authentication layer. Each team receives its own SageMaker AI domain and its own role, while EKS access entries connect those identities to Kubernetes permissions.

Namespaces form the main cluster boundary in the example. Role-based access control is scoped so that a team's identities can interact only with resources in that team's namespace. This makes the boundary explicit at the point where workloads reach EKS instead of relying only on a convention about which resources people should use.

Within a namespace, the guide shows three kinds of HyperPod work: interactive Spaces, distributed PyTorch jobs and inference endpoints. That matters because training, exploration and serving can share the underlying cluster while still arriving through the same team-specific authorization path.

Isolation and fair scheduling solve different problems

A namespace answers where a team may operate. It does not, by itself, decide how a contested pool of accelerators should be divided. AWS places HyperPod Task Governance across the cluster to manage fair resource allocation and scheduling priorities separately from namespace access.

That separation is important for platform teams. Authorization can prevent one group from changing another group's Kubernetes objects, while scheduling governance addresses competition for the same physical capacity. Combining the two controls gives administrators a way to preserve operational boundaries without dedicating an entire cluster to every team.

The guide also keeps observability at the shared platform layer. Teams can use a common cluster while administrators retain a cross-cluster view of workloads and resource use. The evidence describes an architecture pattern, however, not a measured guarantee that every organization will achieve a particular utilization or reliability result.

Storage and cost attribution complete the boundary

Compute separation is only part of a multi-team design. AWS pairs the namespaces with per-team shared directories on Amazon FSx and with per-team or shared Amazon S3 buckets governed by team roles. The intent is to make data access follow the same identity boundaries as workload submission.

For accounting, namespace-level cost allocation associates shared GPU usage with the team that incurred it. AWS presents that visibility as a basis for internal spend reporting and chargeback. It should not be confused with a security control: cost attribution explains who consumed capacity, while IAM and Kubernetes permissions govern who may reach resources.

Together, those layers create a useful operating model. Identity Center establishes the workforce identity, SageMaker AI domains provide team workspaces, EKS access entries and RBAC scope cluster permissions, namespaces contain workloads, Task Governance mediates shared capacity, and storage roles and cost tags extend the separation beyond the scheduler.

What teams still need to verify

This is an AWS-authored reference architecture, not an independent security assessment or a benchmark of real production clusters. Internal confidence is therefore medium. The source establishes how AWS recommends composing its services, but it does not prove that a particular deployment meets an organization's isolation, compliance, performance or budget requirements.

Before adopting the pattern, teams should test their own role mappings, namespace policies, storage permissions, scheduling behavior and cost reports. They should also verify how administrator access, shared observability and any cross-team data paths fit their threat model. Those checks are especially important because logical separation on a shared cluster is a configuration outcome, not a conclusion that follows from using HyperPod alone.

Status

Learning. The architecture and named control layers are documented by AWS; operational outcomes remain deployment-specific and independently unverified.

Sources

Update note: Last reviewed 2026-10-09. We will revise this post if AWS changes the reference architecture or publishes independently testable operational evidence.

Sources

Drafted with AI assistance from source briefs; reviewed for citation completeness and label accuracy.

More Learn coverage