Help improve this page
To contribute to this user guide, choose the Edit this page on GitHub link that is located in the right pane of every page.
Use the NVIDIA DRA driver or device plugin on Amazon EKS
Amazon EKS supports two mechanisms for managing NVIDIA GPU devices in your EKS clusters: the NVIDIA DRA driver for GPUs and the NVIDIA Kubernetes device plugin.
We recommend using the NVIDIA DRA driver for new deployments with Kubernetes versions 1.34 and later when using static capacity provisioning
If you are using GPU sharing features such as multi-instance GPUs (MIG) or time-slicing, we recommend using static configuration with the DRA driver or using the NVIDIA device plugin. Dynamic MIG and time-slicing are in alpha state in the NVIDIA DRA driver. See the NVIDIA DRA driver releases
NVIDIA DRA driver vs. NVIDIA device plugin
| Feature | NVIDIA DRA driver | NVIDIA device plugin |
|---|---|---|
|
Minimum Kubernetes version |
1.34 |
All EKS-supported Kubernetes versions |
|
EKS Compute |
Karpenter (static capacity only), managed node groups, self-managed nodes |
EKS Auto Mode, Karpenter, managed node groups, self-managed nodes |
|
EKS-optimized AMIs |
AL2023 (NVIDIA), Bottlerocket |
AL2023 (NVIDIA), Bottlerocket |
|
Device advertisement |
Rich attributes via |
Integer count of |
|
GPU sharing |
(Alpha) Dynamic MIG, MPS, time-slicing |
(GA) Static MIG, MPS, time-slicing |
|
ComputeDomains |
Manages Multi-Node NVLink (MNNVL) through |
Not supported |
|
Attribute-based selection |
Filter GPUs by model, memory, or other attributes using CEL expressions |
Not supported |
|
Topology-aware EFA allocation |
DRA-native topology-awareness |
Automatic topology-awareness (EKS-optimized AL2023 AMIs only) |
Install the NVIDIA DRA driver
The NVIDIA DRA driver for GPUs manages two types of resources: GPUs and ComputeDomains. It runs two DRA kubelet plugins: gpu-kubelet-plugin and compute-domain-kubelet-plugin. Each can be enabled or disabled separately during installation. This guide focuses on GPU allocation. For using ComputeDomains, see Use P6e-GB200 UltraServers with Amazon EKS.
Prerequisites
-
An Amazon EKS cluster running Kubernetes version 1.34 or later with static capacity provisioned by Karpenter, EKS managed node groups, or self-managed node groups.
-
Nodes with NVIDIA GPU instance types (such as
PorGinstances). -
Nodes with host-level components installed for NVIDIA GPUs. When using the EKS-optimized AL2023 or Bottlerocket NVIDIA AMIs, the host-level NVIDIA driver, CUDA user mode driver, and container toolkit are pre-installed.
-
Helm installed in your command-line environment, see the Setup Helm instructions for more information.
-
kubectlconfigured to communicate with your cluster, see Install or update kubectl for more information.
Procedure
Important
When using the NVIDIA DRA driver for GPU device management, do not deploy it alongside the NVIDIA device plugin on the same node. Doing so can cause silent oversubscription of the underlying devices to multiple pods on the same node.
Disable the built-in NVIDIA device plugin on Bottlerocket
The EKS-optimized Bottlerocket NVIDIA variants include the NVIDIA device plugin and enable it by default. The DRA driver cannot run alongside the device plugin on the same node. Before you use the DRA driver, disable the built-in device plugin on your Bottlerocket GPU nodes. Set settings.kubelet-device-plugins.nvidia.enabled to false in the Bottlerocket node user data.
[settings.kubelet-device-plugins.nvidia] enabled = false
The settings.kubelet-device-plugins.nvidia.enabled setting is available in Bottlerocket version 1.63.0 and later. On earlier versions, the built-in NVIDIA device plugin cannot be disabled. For more information, see bottlerocket-os/bottlerocket pull request #4856
-
Install the NVIDIA DRA driver directly from the Kubernetes SIG OCI registry. To find available versions, see the NVIDIA DRA driver releases
on GitHub. helm install dra-driver-nvidia-gpu \ oci://registry.k8s.io/dra-driver-nvidia/charts/dra-driver-nvidia-gpu \ --version0.4.1\ --create-namespace \ --namespace nvidia \ --set resources.computeDomains.enabled=false \ --set gpuResourcesEnabledOverride=trueFor advanced configuration options, see the NVIDIA DRA driver Helm chart values
on the Kubernetes SIG website. To see the values available for a specific chart version, run helm show values oci://registry.k8s.io/dra-driver-nvidia/charts/dra-driver-nvidia-gpu --version 0.4.1. -
(Optional) To use time-slicing through the DRA driver, add the
TimeSlicingSettingsfeature gate to thehelm installcommand in the previous step. This is an alpha feature that is disabled by default. For more information, see Use GPU time-slicing with the NVIDIA DRA driver.--set featureGates.TimeSlicingSettings=true -
(Optional) To use dynamic MIG through the DRA driver, add the
DynamicMIGfeature gate to thehelm installcommand in the previous step. This is an alpha feature that is disabled by default. You cannot combine theDynamicMIGfeature gate with thePassthroughSupport,NVMLDeviceHealthCheck, orMPSSupportfeature gates. For more information, see Use MIG with the NVIDIA DRA driver.--set featureGates.DynamicMIG=true -
Verify that the DRA driver pods are running.
kubectl get pods -n nvidia -
Verify that the
DeviceClassobjects were created.kubectl get deviceclassNAME AGE gpu.nvidia.com 60s -
Verify that
ResourceSliceobjects are published for your GPU nodes.kubectl get resourcesliceTo request NVIDIA GPUs using the DRA driver, create a
ResourceClaimTemplatethat references thegpu.nvidia.comDeviceClassand reference it in your Pod specification. The following example requests a single GPU. See Topology-aware EFA and GPU/Neuron device allocation for steps to allocate NVIDIA GPUs with topology-aligned EFA interfaces.apiVersion: resource.k8s.io/v1 kind: ResourceClaimTemplate metadata: name: single-gpu spec: spec: devices: requests: - name: gpu exactly: deviceClassName: gpu.nvidia.com count: 1 --- apiVersion: v1 kind: Pod metadata: name: gpu-workload spec: containers: - name: gpu-demo image: public.ecr.aws/amazonlinux/amazonlinux:2023-minimal command: ["/bin/sh", "-c"] args: ["nvidia-smi && tail -f /dev/null"] resources: claims: - name: gpu resourceClaims: - name: gpu resourceClaimTemplateName: single-gpu tolerations: - key: "nvidia.com/gpu" operator: "Exists" effect: "NoSchedule"
Install the NVIDIA Kubernetes device plugin
The NVIDIA Kubernetes device plugin advertises NVIDIA GPUs as nvidia.com/gpu extended resources. You request GPUs in container resource requests and limits.
Prerequisites
-
An Amazon EKS cluster.
-
Nodes with NVIDIA GPU instance types (such as
PorGinstances). -
Nodes with NVIDIA GPU instance types using the EKS-optimized AL2023 NVIDIA AMI. The EKS-optimized Bottlerocket AMIs include the NVIDIA device plugin. You do not need to install it separately.
-
Nodes with host-level components installed for NVIDIA GPUs. When using the EKS-optimized AL2023 or Bottlerocket NVIDIA AMIs, the host-level NVIDIA driver, CUDA user mode driver, and container toolkit are pre-installed.
-
Helm installed in your command-line environment, see the Setup Helm instructions for more information.
-
kubectlconfigured to communicate with your cluster, see Install or update kubectl for more information.
Procedure
-
Add the NVIDIA device plugin Helm chart repository.
helm repo add nvdp https://nvidia.github.io/k8s-device-plugin -
Update your local Helm repository.
helm repo update -
Install the NVIDIA Kubernetes device plugin.
helm install nvdp nvdp/nvidia-device-plugin \ --create-namespace \ --namespace nvidiaOptional: enable GPU Feature Discovery (GFD)
The device plugin advertises
nvidia.com/gpuextended resources and schedules GPU Pods on its own. On the EKS-optimized AL2023 NVIDIA AMI, thenvidia.com/gpu.present=truenode label is already applied at boot bynodeadm, so GPU Feature Discovery (GFD) is not required for basic GPU scheduling.Enable GFD with
--set gfd.enabled=trueif you want the node labeled with detailed GPU attributes—such asnvidia.com/gpu.product,nvidia.com/gpu.memory,nvidia.com/gpu.count, MIG profile labels, and driver/CUDA versions. With these labels, you can target specific GPU types withnodeSelectoror node affinity (for example, scheduling a workload only onto A10G GPUs or onto a particular MIG profile). GPU sharing configurations such as time-slicing and MIG also use these labels. If you don’t need attribute-based node selection or GPU sharing, you can omit the flag.helm install nvdp nvdp/nvidia-device-plugin \ --create-namespace \ --namespace nvidia \ --set gfd.enabled=trueEnable GDRCopy if your workloads use it
k8s-device-pluginv0.19.0 through v0.19.2 enabled the GDRCopy and MOFED features by default. This default enablement was reverted ink8s-device-pluginv0.19.3. As a result of the revert, GDRCopy (gdrdrv) can no longer be enabled per-container with theNVIDIA_GDRCOPY=enabledenvironment variable in a Pod spec, that variable is now ignored.If your workloads use GDRCopy (GPUDirect RDMA copy), you must enable it on the device plugin at install time by setting
gdrcopyEnabled=true:helm upgrade --install nvdp nvdp/nvidia-device-plugin \ --namespace nvidia \ --create-namespace \ --set gdrcopyEnabled=true(Add
--set gfd.enabled=trueas well if you also want GPU Feature Discovery labels, as described previously.)If you manage the NVIDIA device plugin through the NVIDIA GPU Operator
on GitHub, the operator dynamically sets GDRCOPY_ENABLED=truewhen thegdrdrvkernel module is loaded on the node.For more information, see NVIDIA k8s-device-plugin issue #1692
on GitHub. Note
You can also install and manage the NVIDIA Kubernetes device plugin using the NVIDIA GPU Operator
on GitHub, which automates the management of all NVIDIA software components needed to provision GPUs. -
Verify the NVIDIA device plugin DaemonSet is running.
kubectl get ds -n nvidia nvdp-nvidia-device-pluginNAME DESIRED CURRENT READY UP-TO-DATE AVAILABLE NODE SELECTOR AGE nvdp-nvidia-device-plugin 2 2 2 2 2 <none> 60s -
Verify that your nodes have allocatable GPUs.
kubectl get nodes "-o=custom-columns=NAME:.metadata.name,GPU:.status.allocatable.nvidia\.com/gpu"An example output is as follows.
NAME GPU ip-192-168-11-225.us-west-2.compute.internal 1 ip-192-168-24-96.us-west-2.compute.internal 1
Request NVIDIA GPUs in a Pod
To request NVIDIA GPUs using the device plugin, specify the nvidia.com/gpu resource in your container resource requests and limits.
apiVersion: v1 kind: Pod metadata: name: nvidia-smi spec: restartPolicy: OnFailure containers: - name: gpu-demo image: public.ecr.aws/amazonlinux/amazonlinux:2023-minimal command: ["/bin/sh", "-c"] args: ["nvidia-smi && tail -f /dev/null"] resources: limits: nvidia.com/gpu: 1 requests: nvidia.com/gpu: 1 tolerations: - key: "nvidia.com/gpu" operator: "Equal" value: "true" effect: "NoSchedule"
To run this test, apply the manifest and view the logs:
kubectl apply -f nvidia-smi.yaml kubectl logs nvidia-smi
An example output is as follows.
+-----------------------------------------------------------------------------------------+ | NVIDIA-SMI XXX.XXX.XX Driver Version: XXX.XXX.XX CUDA Version: XX.X | |-----------------------------------------+------------------------+----------------------+ | GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC | | Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. | | | | MIG M. | |=========================================+========================+======================| | 0 NVIDIA L4 On | 00000000:31:00.0 Off | 0 | | N/A 27C P8 11W / 72W | 0MiB / 23034MiB | 0% Default | | | | N/A | +-----------------------------------------+------------------------+----------------------+ +-----------------------------------------------------------------------------------------+ | Processes: | | GPU GI CI PID Type Process name GPU Memory | | ID ID Usage | |=========================================================================================| | No running processes found | +-----------------------------------------------------------------------------------------+