Setting up task governance for Ray - Amazon SageMaker AI
Services or capabilities described in AWS documentation might vary by Region. To see the differences applicable to the AWS European Sovereign Cloud Region, see the AWS European Sovereign Cloud User Guide.

Setting up task governance for Ray

Set up Task Governance on your cluster first, then complete the Ray-specific steps on this page. For the concepts and the console setup, see Setup for SageMaker HyperPod task governance.

Namespaces that need a compute allocation

Task Governance admits a Ray workload only in a namespace that has a compute allocation. Create an allocation for every namespace where you create Ray workloads. A workload in a namespace with no allocation stays pending and is never admitted. For the console steps, see Policies.

Gang scheduling

Confirm that gang scheduling is enabled for your cluster. A Ray cluster needs its head and all of its workers running together, so without gang scheduling a partially scheduled cluster holds capacity without making progress. Task Governance implements gang scheduling with the Kueue waitForPodsReady feature, which evicts and requeues a workload whose pods do not all become ready within the configured timeout. For the configuration settings, see Using gang scheduling in Amazon SageMaker HyperPod task governance.

Verify

Create a small RayCluster in an allocated namespace and confirm it reaches a running state. If it stays pending, confirm the namespace has a compute allocation with room for the declared cluster size.