GPU scheduling on Kubernetes with Dynamic Resource Allocation (DRA): the 2026 guide

GPU scheduling on Kubernetes with Dynamic Resource Allocation (DRA): the 2026 guide

GPU scheduling in Kubernetes used to be deceptively simple: install the NVIDIA device plugin, request nvidia.com/gpu: 1, and let the default scheduler find a node with one available GPU. That model got many clusters into production, but it encoded the wrong abstraction. A modern GPU is not just an integer. It has memory size, architecture, interconnect, topology, partitioning modes, sharing modes, and health state.

Related reading: Cluster Autoscaler vs Karpenter.

Dynamic Resource Allocation (DRA) is Kubernetes’ answer to that mismatch. The core DRA APIs graduated to GA in Kubernetes 1.34, with the stable resource.k8s.io/v1 API enabled by default. In 2026, this matters because the ecosystem around AI infrastructure has also moved: NVIDIA donated its DRA Driver for GPUs to the Kubernetes community under CNCF governance at KubeCon Europe 2026, Kueue has native concepts for DRA-aware quota management, KAI Scheduler is a CNCF Sandbox project for large GPU fleets, and inference stacks such as vLLM and KServe are becoming the runtime layer above the scheduler.

This is not a “replace one YAML key with another” migration. DRA changes where device knowledge lives. Instead of asking Kubernetes for a count of opaque extended resources, workloads request a claim against a class of devices, and the scheduler allocates a concrete device that satisfies the claim.

Why the device-plugin model is reaching its limits

The Kubernetes device plugin framework exposes vendor devices to the kubelet. For NVIDIA GPUs, the traditional resource name is nvidia.com/gpu, requested in container resources.requests and resources.limits. This remains useful, especially for simple clusters.

The limitation is in the API shape. Kubernetes extended resources are integer resources and cannot be overcommitted. The Kubernetes documentation also states that devices cannot be shared between containers through the basic extended-resource model. That is fine for “one pod owns one whole GPU”. It is much weaker for LLM inference, mixed training queues, fractional capacity, MIG profiles, topology-sensitive multi-GPU jobs, and heterogeneous node pools.

The device-plugin model also pushes too much meaning into out-of-band policy. If you need A100s rather than L4s, you usually add node labels, node affinity, taints, or separate node groups. If you need a MIG slice, you configure GPU Operator and device plugin strategy, then expose separate resource names or labels. If you need low-latency multi-GPU placement, you combine scheduler plugins, topology labels, and workload-specific conventions.

Those workarounds fragment the source of truth. The scheduler sees integer capacity. The driver knows device details. The autoscaler knows node templates. The ML platform knows model requirements. DRA gives Kubernetes a structured device model so those systems can coordinate through API objects.

For node provisioning, this does not remove the need for good autoscaling. You still need a node autoscaler that can bring up GPU capacity with the right labels, taints, AMI, driver stack, and instance family. If you run on AWS, /cluster-autoscaler-vs-karpenter/ is directly relevant: GPU pods sitting Pending are often a node provisioning problem before they are a scheduler problem. On EKS, /eks-auto-mode/ is also worth reading because managed node lifecycle and accelerator support change how much of the stack you own.

What DRA actually adds

DRA introduces Kubernetes APIs for claiming devices. The stable API group is resource.k8s.io/v1. The important objects are:

API objectScopePurpose
DeviceClassclusterDefines a category of devices and optional selectors/configuration. Claims reference a DeviceClass.
ResourceSliceclusterPublished by DRA drivers. Describes available devices, attributes, capacity, and node access.
ResourceClaimnamespaceRequests access to devices. The scheduler allocates concrete devices into the claim status.
ResourceClaimTemplatenamespaceTemplate for per-pod ResourceClaims, similar in spirit to volume claim templates.
Pod spec.resourceClaimspodMakes a ResourceClaim or ResourceClaimTemplate available to the pod.
Container resources.claimscontainerAttaches a named claim to a specific container.

The control flow is simple. A DRA driver publishes devices as ResourceSlice objects. A cluster administrator or driver provides DeviceClass objects. A workload creates a ResourceClaim or references a ResourceClaimTemplate. During scheduling, Kubernetes evaluates the claim against ResourceSlices, picks devices on nodes where the pod can run, stores the result in ResourceClaim status, and the driver prepares the device for the pod.

This is close to the PersistentVolumeClaim mental model: a pod references a claim with requirements, and Kubernetes binds it to something concrete. With DRA, the pod claims a device from a class, with selectors and constraints that drivers and the scheduler can reason about.

Do not overstate this: DRA is not live GPU hot-plugging, and it is not a model-serving platform. It is a scheduling-time allocation framework. You still need CPU and memory requests to be correct, as covered in /2026-05-kubernetes-resource-requests-limits/. You still need queueing for batch fairness and an inference runtime for serving.

A minimal DRA GPU request

The exact DeviceClass names and available attributes depend on the installed driver. NVIDIA’s NIM Operator documentation says the NVIDIA DRA driver deploys a default gpu.nvidia.com DeviceClass for physical GPUs. Always verify this in your cluster:

kubectl get deviceclasses
kubectl get resourceslices

For a single GPU per pod, use a ResourceClaimTemplate:

apiVersion: resource.k8s.io/v1
kind: ResourceClaimTemplate
metadata:
  name: one-nvidia-gpu
  namespace: ai
spec:
  spec:
    devices:
      requests:
      - name: gpu
        exactly:
          deviceClassName: gpu.nvidia.com
          allocationMode: ExactCount
          count: 1
---
apiVersion: batch/v1
kind: Job
metadata:
  name: dra-gpu-smoke-test
  namespace: ai
spec:
  completions: 1
  parallelism: 1
  template:
    spec:
      restartPolicy: Never
      resourceClaims:
      - name: gpu
        resourceClaimTemplateName: one-nvidia-gpu
      containers:
      - name: cuda
        image: nvcr.io/nvidia/cuda:12.5.1-base-ubuntu22.04
        command: ["nvidia-smi"]
        resources:
          claims:
          - name: gpu
          requests:
            cpu: "1"
            memory: 1Gi
          limits:
            memory: 1Gi

This manifest uses the stable resource.k8s.io/v1 API and the pod-level resourceClaims plus container-level resources.claims fields shown in the Kubernetes DRA task documentation. It deliberately does not use the legacy nvidia.com/gpu resource request.

In production, you will usually add selectors. Kubernetes supports CEL selectors in DRA claims, but attribute names are driver-specific. Inspect ResourceSlices before standardizing selectors:

kubectl get resourceslices -o yaml

For example, a platform team might publish DeviceClasses such as “inference-l4”, “training-h100”, or “mig-1g-10gb” instead of asking application teams to write low-level CEL expressions.

MIG, sharing, and topology awareness

MIG is where DRA becomes more than nicer syntax. With the device-plugin model, MIG support works, but the cluster often ends up with a mixture of resource names, node labels, and operational conventions. DRA lets the driver publish device shapes and capacities in ResourceSlices.

NVIDIA’s current DRA driver documentation is cautious: the driver manages GPUs and ComputeDomains; ComputeDomains are officially supported for robust and secure Multi-Node NVLink, while some GPU allocation features are still described as exploratory in the upstream README. NVIDIA’s NIM Operator documentation already covers Kubernetes v1.34 or later, resource.k8s.io/v1, the gpu.nvidia.com DeviceClass, full GPU versus MIG decisions, and the GPU Operator path. Pin driver versions and test the exact mode you intend to offer.

Topology matters at two levels.

First, there is node-local topology: NUMA locality, PCIe lanes, NVLink, and whether devices are close to the CPUs and memory the pod uses. Kubernetes has long had a Topology Manager in kubelet, but DRA gives the scheduler more structured information before placement.

Second, there is fleet topology: racks, blocks, zones, and inter-node fabric. Kueue’s Topology Aware Scheduling documentation targets AI/ML workloads where pod-to-pod bandwidth affects runtime and cost. DRA handles device allocation; Kueue decides when a workload should be admitted and how scarce quota should be shared.

Preemption belongs in that same layered view. Kubernetes scheduler preemption and Kueue preemption can make room for higher-priority workloads, but DRA itself is the device allocation substrate. In practice, you combine DRA claims, PriorityClasses, Kueue ClusterQueues, and possibly KAI Scheduler policies to get predictable multi-tenant behavior.

What NVIDIA’s CNCF donation changes

On March 24, 2026, NVIDIA announced at KubeCon Europe in Amsterdam that it was donating the NVIDIA DRA Driver for GPUs to CNCF, moving it from vendor governance to community ownership under the Kubernetes project.

It changes the risk profile for platform teams. A GPU DRA driver under Kubernetes community governance is easier to treat as part of the cloud-native substrate, and it creates a clearer collaboration point for cloud providers, Kubernetes SIGs, Kueue, KAI Scheduler, and inference platforms.

It does not make NVIDIA-specific hardware vendor-neutral. CUDA, MIG, NVLink, GPU Operator, and driver lifecycle remain NVIDIA concerns. What improves is the Kubernetes integration surface for advertising, claiming, allocating, and preparing devices.

NVIDIA also announced that KAI Scheduler had been onboarded as a CNCF Sandbox project. CNCF describes KAI Scheduler as a Kubernetes scheduler for optimizing GPU resource allocation for AI workloads in large-scale clusters. DRA models and allocates devices; KAI provides AI-focused scheduling policy; Kueue provides quota, admission, fair sharing, and preemption; KServe and vLLM provide the inference serving layer.

For production inference, vLLM and KServe are consumers of GPU scheduling, not replacements for it. vLLM documents KServe integration for distributed model serving, and KServe documents multi-node, multi-GPU inference using a vLLM serving runtime.

Migration: device plugin to DRA

Do not migrate the whole fleet in one step. A practical migration looks like this:

  1. Inventory existing GPU workloads. Classify them by device shape: whole GPU training, small inference, MIG-friendly inference, multi-GPU single-node, and multi-node training.
  2. Upgrade the control plane and nodes to Kubernetes 1.34 or later. Verify resource.k8s.io/v1 is available with kubectl api-resources | grep resource.k8s.io.
  3. Install or upgrade the GPU Operator and NVIDIA DRA driver in a small, isolated GPU node pool.
  4. Verify DeviceClass and ResourceSlice objects before writing application selectors.
  5. Define platform-owned DeviceClasses for common use cases. Prefer “h100-training”, “l4-inference”, or “mig-small” over asking every team to understand device internals.
  6. Convert one non-critical workload from nvidia.com/gpu to a ResourceClaimTemplate. Keep CPU and memory requests unchanged unless you are intentionally resizing the workload.
  7. Add Kueue for batch admission if multiple teams compete for GPUs. Use ClusterQueues, ResourceFlavors, cohorts, fair sharing, and preemption policies.
  8. Validate node provisioning. If the claim is valid but no node exists, the autoscaler still has to create the right GPU node.
  9. Roll into serving platforms. For KServe/vLLM, verify how the serving controller passes or creates DRA claims.
  10. Retire legacy paths only after observability, rollback, and quota policies are in place.

During migration, it is reasonable to run legacy device-plugin workloads and DRA workloads side by side on separate node pools. Avoid advertising the same physical GPU through two allocation systems to the same scheduling domain unless the driver documentation explicitly supports that configuration.

Device plugin vs DRA

CapabilityDevice plugin / extended resourceDRA
Workload requestnvidia.com/gpu: 1 integer resourceResourceClaim or ResourceClaimTemplate
API maturityDevice plugin framework exists since Kubernetes 1.10 betaCore DRA APIs GA in Kubernetes 1.34
Device attributesMostly external labels and conventionsStructured attributes/capacity through ResourceSlices
SharingLimited by extended-resource model; vendor-specific sharing modesClaims can model sharing patterns when supported by driver and feature set
MIG / partitionsWorks through vendor configuration and resource exposureBetter fit for requesting specific device shapes
Scheduler awarenessCounts resources on nodesAllocates concrete devices during scheduling
Autoscaler visibilityOften depends on node templates and resource namesStructured parameters improve simulation potential, but autoscaler support still matters
Best fitSimple whole-GPU workloadsHeterogeneous, partitioned, shared, or topology-sensitive GPU fleets

When you do not need DRA yet

DRA is not mandatory for every GPU cluster in 2026.

If every workload needs exactly one whole GPU on a homogeneous node pool, the device plugin model may be simpler. If your managed Kubernetes provider does not support the driver path you need, waiting is rational. If your bottleneck is cold node provisioning, image pull time, model download time, or bad CPU/memory requests, DRA will not fix that by itself.

Pilot DRA when multiple teams compete for expensive GPUs, you run mixed GPU models, MIG or sharing is a first-class requirement, multi-GPU topology affects performance, or you need cleaner integration between scheduling, quota, and AI workload platforms.

FAQ

Is DRA GA in Kubernetes?

The core DRA APIs graduated to GA in Kubernetes 1.34, using resource.k8s.io/v1 and enabled by default. Some surrounding features continued to mature later; for example, Kubernetes documentation in 2026 marks certain DRA task flows as v1.35 [stable] and DRA prioritized lists as v1.36 [stable].

Does DRA replace the NVIDIA device plugin?

For DRA workloads, the claim replaces the nvidia.com/gpu extended-resource request. Operationally, many clusters will run both models during migration. The NVIDIA DRA driver is the relevant component for DRA-based GPU allocation.

Can I request fractional GPUs with DRA?

DRA provides a framework for richer requests and sharing, but the exact behavior depends on Kubernetes feature maturity and the driver. NVIDIA documents MIG and time-slicing paths in its GPU stack, while DRA consumable capacity adds a Kubernetes model for sharing device capacity. Validate the exact mode you want before promising fractional GPU self-service.

Is DRA enough for multi-tenant GPU scheduling?

No. DRA allocates devices. For fair sharing, quota, admission, and preemption across teams, add Kueue or a scheduler layer such as KAI Scheduler, depending on your workload shape.

Does DRA help with inference?

Yes, but indirectly. DRA makes the GPU request more accurate. vLLM and KServe still handle serving concerns such as model runtime, scaling, routing, and multi-node inference patterns.

CTA: pilot DRA in one GPU node pool

The right first step is one GPU node pool. Install the NVIDIA DRA driver, verify DeviceClass and ResourceSlice objects, and convert one smoke-test Job plus one real low-risk workload to ResourceClaimTemplate. Measure scheduling latency, claim allocation status, GPU utilization, failure modes, and autoscaler behavior. If the pilot is boring, expand it to one tenant queue. If it is noisy, fix the platform contract first.

Sources

  • Kubernetes v1.34 DRA GA announcement: https://kubernetes.io/blog/2025/09/01/kubernetes-v1-34-dra-updates/
  • Kubernetes Dynamic Resource Allocation concepts: https://kubernetes.io/docs/concepts/scheduling-eviction/dynamic-resource-allocation/
  • Kubernetes DRA workload task and YAML fields: https://kubernetes.io/docs/tasks/configure-pod-container/assign-resources/allocate-devices-dra/
  • Kubernetes ResourceClaim API reference: https://kubernetes.io/docs/reference/kubernetes-api/resource/resource-claim-v1/
  • Kubernetes ResourceSlice API reference: https://kubernetes.io/docs/reference/kubernetes-api/resource/resource-slice-v1/
  • Kubernetes device plugin documentation: https://kubernetes.io/docs/concepts/extend-kubernetes/compute-storage-net/device-plugins/
  • Kubernetes DRA consumable capacity: https://kubernetes.io/blog/2025/09/18/kubernetes-v1-34-dra-consumable-capacity/
  • KEP-4381 DRA structured parameters: https://github.com/kubernetes/enhancements/tree/master/keps/sig-node/4381-dra-structured-parameters
  • NVIDIA DRA Driver for GPUs repository: https://github.com/kubernetes-sigs/dra-driver-nvidia-gpu
  • NVIDIA KubeCon Europe 2026 DRA driver donation announcement: https://blogs.nvidia.com/blog/nvidia-at-kubecon-2026/
  • NVIDIA NIM Operator DRA support: https://docs.nvidia.com/nim-operator/latest/dra.html
  • NVIDIA GPU Operator sharing documentation: https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/gpu-sharing.html
  • CNCF KAI Scheduler project page: https://www.cncf.io/projects/kai-scheduler/
  • Kueue overview and DRA concepts: https://kueue.sigs.k8s.io/docs/overview/
  • Kueue Topology Aware Scheduling: https://kueue.sigs.k8s.io/docs/concepts/topology_aware_scheduling/
  • vLLM KServe integration: https://docs.vllm.ai/en/stable/deployment/integrations/kserve/
  • KServe multi-node/multi-GPU vLLM inference: https://kserve.github.io/website/docs/model-serving/generative-inference/multi-node

EKS Auto Mode: What It Actually Changes (and What It Doesn’t)

EKS Auto Mode: What It Actually Changes (and What It Doesn't)

What EKS Auto Mode is

EKS Auto Mode, generally available on December 1, 2024, shifts more Kubernetes infrastructure responsibility from you to AWS. You still run an EKS cluster in your AWS account, and your workloads still use the Kubernetes API, but AWS takes over much of the compute, storage, networking, node lifecycle, and core add-on management that platform teams usually wire together themselves.

For compute, Auto Mode uses Karpenter-style provisioning under the hood. When pods are unschedulable, Auto Mode provisions nodes that fit the workload’s requirements: instance family, size, architecture, capacity type, and availability zone. When capacity is no longer useful, it can consolidate and terminate nodes.

The important framing is this: Auto Mode is not “EKS without nodes.” It is EKS where AWS manages the node lifecycle more aggressively. You own the workloads, their scheduling requirements, their disruption behavior, and the operational consequences of those choices. AWS owns more of the infrastructure plumbing.


What it replaces

Before Auto Mode, running EKS in production usually meant choosing and operating several layers yourself:

Managed node groups: You chose instance types, defined scaling ranges, managed AMI updates, handled node draining, and configured Cluster Autoscaler or another scaling mechanism.

Self-managed Karpenter: More flexible than managed node groups, but you owned the Karpenter controller, IAM, NodePools, EC2NodeClasses, disruption settings, upgrades, and failure modes.

Fargate: AWS-managed compute per pod, with no node management, but no DaemonSets, a narrower workload compatibility envelope, and a different cost model.

EKS Auto Mode replaces a large part of that platform assembly with a managed model: declare workload intent and high-level compute constraints; AWS provisions and manages the EC2 instances behind it.


How it works in practice

You create or update an EKS cluster with Auto Mode enabled. The default setup can use AWS-managed built-in node pools. If you need more control, you create a NodeClass for Auto Mode infrastructure settings and a Karpenter NodePool for workload-facing scheduling constraints.

# NodeClass: EKS Auto Mode infrastructure settings for managed EC2 nodes.
apiVersion: eks.amazonaws.com/v1
kind: NodeClass
metadata:
  name: private-compute
spec:
  subnetSelectorTerms:
    - tags:
        kubernetes.io/role/internal-elb: "1"
  securityGroupSelectorTerms:
    - tags:
        aws:eks:cluster-name: prod-eks
  ephemeralStorage:
    size: "100Gi"
# NodePool: workload-facing constraints for nodes that Auto Mode may provision.
apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
  name: general-purpose
spec:
  template:
    spec:
      nodeClassRef:
        group: eks.amazonaws.com
        kind: NodeClass
        name: private-compute
      requirements:
        - key: karpenter.sh/capacity-type
          operator: In
          values: ["on-demand", "spot"]
        - key: kubernetes.io/arch
          operator: In
          values: ["amd64", "arm64"]
        - key: eks.amazonaws.com/instance-category
          operator: In
          values: ["c", "m", "r"]
  limits:
    cpu: "1000"
    memory: 1000Gi
  disruption:
    consolidationPolicy: WhenEmptyOrUnderutilized
    consolidateAfter: 1m

Those API groups are the Auto Mode-specific split documented by AWS: apiVersion: eks.amazonaws.com/v1, kind: NodeClass for Auto Mode node infrastructure, and apiVersion: karpenter.sh/v1, kind: NodePool for scheduling and capacity constraints.

The NodeClass is where you express AWS infrastructure placement and node-level defaults. The NodePool is where you express what kind of capacity is acceptable for workloads. Do not copy a self-managed Karpenter EC2NodeClass into Auto Mode; Auto Mode uses its own NodeClass API.

Auto Mode provisions nodes when pods are pending, consolidates when nodes are underutilized, and replaces nodes during maintenance or scale-down. AMI and node lifecycle updates are handled by AWS. AWS says Auto Mode AMIs are generally released weekly with CVE and security fixes, and Auto Mode nodes have a maximum lifetime of 21 days, which you can reduce. Your application still has to tolerate the disruption: a bad PodDisruptionBudget, strict affinity rule, or singleton stateful workload can still block or degrade a replacement.

Built-in components are managed differently than in a classic EKS build. AWS lists pod networking, service networking, cluster DNS, autoscaling, block storage, load balancer controller, Pod Identity agent, and node monitoring agent as Auto Mode capabilities. With Auto Mode compute, common add-ons such as Amazon VPC CNI, kube-proxy, CoreDNS, Amazon EBS CSI Driver, and EKS Pod Identity Agent become redundant for Auto Mode nodes, and the relevant controllers can run on AWS-owned infrastructure rather than as visible pods in your account. You can still install AWS Load Balancer Controller in an Auto Mode cluster when you need both models during migration, but AWS does not support directly migrating existing load balancers from AWS Load Balancer Controller to Auto Mode; use IngressClass or loadBalancerClass boundaries and plan blue-green migration. Treat this as a change in ownership, not as a reason to skip validation.

The workload support matrix is broader than Fargate, but not identical to self-managed nodes:

CapabilityAuto Mode status
EC2 SpotSupported through karpenter.sh/capacity-type requirements such as spot, on-demand, and reserved
Graviton / arm64Supported through kubernetes.io/arch: arm64 and supported Graviton instance families
GPU / acceleratorsSupported for documented accelerated families; Auto Mode manages NVIDIA, Trainium, and Inferentia drivers/device plugins for supported instance types
Windows nodesNot supported
DaemonSetsSupported as Kubernetes DaemonSets, but host-level assumptions must be validated against locked-down managed nodes

What you gain

Reduced operational surface. Node group management, AMI lifecycle, Cluster Autoscaler tuning, Karpenter controller upgrades, and a chunk of add-on wiring move out of your day-to-day scope.

Better provisioning shape by default. Dynamic provisioning is usually a better fit than fixed node group shapes. You get nodes that more closely match actual pod requirements instead of trying to pre-plan a small set of instance types.

Automatic node patching. AWS manages the node image and replacement flow. That reduces toil, but it also means your workloads need disruption policies that let AWS replace nodes safely.

Faster cluster bootstrapping. A new Auto Mode cluster can get to a usable production baseline faster than a hand-assembled EKS cluster with node groups, autoscaling, networking add-ons, storage drivers, and load balancer controllers.

Native Spot integration. Auto Mode can use Spot capacity through NodePool requirements, but you still need workload-level interruption tolerance: replicas, budgets, graceful shutdown, and queue semantics where relevant.


What you give up

Node-level access. Auto Mode nodes are intentionally locked down compared with traditional self-managed nodes. If your incident response process assumes SSH, SSM, manual package inspection, or ad hoc host changes, it needs to change.

Custom AMIs. You cannot treat the node image as your own artifact. AWS determines the operating system and AMI for Auto Mode managed instances; you cannot directly access the instance or install software on it. If your organization requires internally built, hardened, or certified AMIs, Auto Mode is likely blocked.

Unrestricted host agents. Kubernetes DaemonSets are supported, but they are the sharp edge. Some node agents work; others do not. Anything that assumes privileged host access, custom kernel modules, hostPath writes, IMDS access without hostNetwork, or low-level runtime integration needs a proof of compatibility.

Less tuning surface. You give up direct control over kubelet flags, container runtime configuration, bootstrap scripts, and arbitrary node setup. That is the point of the product, but it is also the boundary.

Different cost visibility. Managed node groups make capacity easier to reason about because you chose it up front. Auto Mode changes capacity dynamically, so cost control moves toward budgets, labels, reports, and workload-level resource hygiene.


EKS Auto Mode vs managed node groups vs Fargate

Auto ModeManaged Node GroupsFargate
Node managementAWS manages node lifecycleShared: AWS manages the group primitive, you manage capacity shape and many updatesAWS-managed per-pod compute
AMI updatesAutomatic through AWS-managed node replacementYou schedule and operate rolling updatesN/A
Instance selectionDynamic through NodePoolsYou choose instance types and scaling rangesNot exposed
Custom AMIsNo; AWS determines the AMIYesNo
DaemonSetsSupported, but validate host-access assumptionsYesNo
SSH / node accessRestrictedUsually available if you enable itNo
Spot supportYes, through capacity-type requirementsYes, with node group or Karpenter designNo; Amazon EKS does not support Fargate Spot
Cost modelEKS control plane + EC2 + EKS Auto Mode feeEKS control plane + EC2EKS control plane + Fargate pod pricing
Operational burdenLow for nodes, medium for workload compatibilityMediumLow for nodes, medium for compatibility
Right forDefault candidate for teams that do not need node customizationRegulated/custom node environments and mature platform teamsWorkloads that fit Fargate’s restrictions and want per-pod isolation
Hidden constraints / gotchasAWS controls node image and lifecycle; DaemonSets, privileged pods, hostPath, PDBs, topology rules, and node agents can block migrationYou still own AMI drift, autoscaler tuning, disruption handling, and capacity fragmentationNo DaemonSets, limited host-level integrations, different networking/storage constraints, and less flexibility for mixed workload shapes

The Auto Mode pricing

Do not model Auto Mode as “EC2 plus a generic percentage” unless you have pulled the actual rate for your region and instance mix. The official structure is:

Total EKS Auto Mode cluster cost =
  EKS control plane cost
  + normal EC2 cost for instances launched and managed by Auto Mode
  + EKS Auto Mode management fee on those managed EC2 instances
  + normal surrounding AWS costs: EBS, load balancers, data transfer, CloudWatch, etc.

EKS Auto Mode management fee =
  sum of Auto Mode-managed instance runtime
  x the regional EKS Auto Mode management rate for each EC2 instance type

The EKS Auto Mode fee is applied to the EC2 instances that Auto Mode launches and manages. It is billed in addition to the normal EC2 charge and in addition to the EKS control plane charge. AWS bills the Auto Mode fee per second with a one-minute minimum, and the charge is independent of whether the underlying EC2 capacity is On-Demand, Spot, covered by Reserved Instances, or covered by Compute Savings Plans.

Do not treat the fee as a contractual flat percentage. The official pricing page describes it as a management fee that varies by EC2 instance type, and AWS pricing data is regional. The public pricing example for US West (Oregon) shows c6a.2xlarge at $0.306/hour for EC2 plus $0.03672/hour for Auto Mode, c6a.4xlarge at $0.612/hour plus $0.07344/hour, m5a.2xlarge at $0.344/hour plus $0.04128/hour, and m5a.xlarge at $0.172/hour plus $0.02064/hour. Those examples equal 12% of the listed On-Demand EC2 rate, but the safe formula for real planning is: sum(instance-hours by instance type and region x published Auto Mode management rate).

The practical cost question is not “is there a premium?” There is. The useful question is whether the premium is lower than the engineering time, incident risk, and opportunity cost of operating node lifecycle yourself.

For small teams, the answer may be yes even if the raw bill increases. For high-scale, cost-sensitive platforms, the answer needs real data: compare current EC2 waste, bin-packing efficiency, Spot usage, interruption rate, and platform maintenance time against an Auto Mode pilot.


When to use EKS Auto Mode

Use Auto Mode if:

  • You run EKS on AWS and do not have a hard requirement to manage nodes yourself
  • You do not require custom AMIs or custom node bootstrap logic
  • You want to reduce the operations surface for node lifecycle management
  • You want Karpenter-like provisioning without operating Karpenter yourself
  • Your workloads are mostly stateless or disruption-tolerant
  • Your observability, security, and storage agents are compatible with Auto Mode

Stick with managed node groups if:

  • Your organization requires internally certified or hardened AMIs
  • You need specific kernel configuration, kubelet flags, bootstrap scripts, or host packages
  • You depend on privileged DaemonSets or host-level security tooling that Auto Mode cannot support
  • You are in a regulated environment where the node image supply chain must be owned internally
  • Your platform team already operates Karpenter well and values the extra control

Use Fargate if:

  • You specifically want per-pod compute isolation
  • Your workload does not need DaemonSets or host-level integrations
  • You accept Fargate’s scheduling, networking, storage, and observability constraints
  • You want to avoid managing EC2 capacity entirely for a narrow class of workloads

Migration from managed node groups

Migrating an existing cluster to Auto Mode is supported, but it is not a one-command operational migration. AWS supports enabling Auto Mode on existing clusters, but you must update the cluster IAM role permissions and trust policy, enable compute, block storage, and load balancing capabilities together, and meet required add-on versions when those add-ons are installed. AWS also calls out unsupported direct migrations for EBS volumes from the standard EBS CSI provisioner to the Auto Mode EBS CSI provisioner, existing load balancers from AWS Load Balancer Controller to Auto Mode, and clusters using alternative CNIs or other unsupported networking configurations. A conservative path looks like this:

  1. Enable Auto Mode on a non-production cluster running Kubernetes 1.29 or greater.
  2. Inventory workloads by scheduling assumptions: node selectors, affinities, tolerations, topology spread, PDBs, privileged mode, hostPath, local storage, and DaemonSet dependencies.
  3. Create or select the relevant Auto Mode NodeClass and NodePool resources.
  4. Move a low-risk namespace first by changing selectors, tolerations, or labels so pods land on Auto Mode nodes.
  5. Watch scheduling, replacement, load balancer behavior, persistent volume provisioning, logging, metrics, and security events.
  6. Taint old node groups to stop new scheduling once the pilot workloads are stable.
  7. Drain old nodes gradually and delete old managed node groups only after workload owners have signed off.

The hardest part is usually not enabling Auto Mode. It is discovering which workloads and platform agents quietly depended on a mutable node.


What breaks when you migrate

Auto Mode changes the node contract. The Kubernetes API still looks familiar, but the host underneath is no longer yours in the same way.

DaemonSets need a compatibility audit. Logging agents, metrics agents, service mesh node components, security scanners, CSI node plugins, and custom infrastructure daemons often assume host access. Datadog, Falco, custom CSI drivers, eBPF agents, file integrity tools, and in-house node agents should be tested explicitly rather than assumed compatible.

PodDisruptionBudgets can block AWS-managed maintenance. If every critical Deployment has maxUnavailable: 0, or singleton workloads have no safe disruption path, node replacement becomes harder. Auto Mode can manage nodes, but it cannot make an application disruption-tolerant after the fact.

nodeSelector and affinity rules can strand pods. Workloads pinned to old node group labels, instance types, capacity labels, zones, or custom AMI labels may never schedule on Auto Mode capacity. Replace legacy labels with stable requirements that Auto Mode can satisfy.

topologySpreadConstraints can become too strict. Auto Mode provisions capacity dynamically, but strict zone spreading plus narrow selectors can create unschedulable pods. Check whenUnsatisfiable, label selectors, and minimum domain assumptions.

Privileged pods and hostPath volumes are migration blockers until proven otherwise. Anything that needs /var/lib, /proc, /sys, container runtime sockets, kernel capabilities, or host networking deserves a separate test. Some patterns are fundamentally at odds with locked-down managed nodes.

Observability and security agents may lose host assumptions. Agents that expect direct node access, host package installation, kernel modules, eBPF privileges, or container runtime socket access can fail partially. The dangerous failure mode is not “pod CrashLoopBackOff”; it is silent loss of telemetry or enforcement.

Storage drivers must be reviewed. EBS integration is part of Auto Mode, but it uses the Auto Mode EBS CSI provisioner ebs.csi.eks.amazonaws.com, not the standard EBS CSI provisioner ebs.csi.aws.com. Custom CSI drivers, EFS patterns, snapshot controllers, and topology-aware storage classes should be validated. Pay particular attention to provisioner names, volume binding mode, encryption settings, and IAM assumptions.

Runbooks need rewriting. “SSH to the node and inspect X” is not a valid first response anymore. Incident procedures should move toward kubectl describe, events, logs, ephemeral debug containers where supported, cloud-side metrics, and vendor-supported diagnostics.


The honest assessment

EKS Auto Mode is a good default candidate for many teams running Kubernetes on AWS. The operational simplification is real: node provisioning, AMI updates, core add-on integration, and scaling behavior are areas where teams burn time and create incidents.

The constraints are also real. Custom AMIs, unrestricted host access from DaemonSets, privileged pods, custom CSI drivers, and strict disruption policies are the common blockers. If your platform depends on those, Auto Mode is not a free upgrade.

For teams without those constraints, Auto Mode should be evaluated early for new EKS clusters. For existing clusters, it should be treated as a migration project, not a checkbox. The right question is not whether Auto Mode is “better” than managed node groups. The right question is which operational contract your workloads can actually live with.

The pattern is the same as with managed infrastructure generally: the more your organization can treat nodes as replaceable capacity, the more value you get. The more your platform treats nodes as customized machines, the less Auto Mode fits.


1-week pilot: evaluate Auto Mode without risking production

Use a short pilot to answer compatibility and economics questions before touching production.

  1. Create a test cluster or clone a representative non-production cluster. Use the same region, Kubernetes minor version, VPC shape, IAM model, ingress pattern, and storage classes where possible.
  2. Enable Auto Mode and deploy one custom NodeClass and NodePool. Keep the first pool boring: on-demand capacity, two or three common instance families, and the same private subnet pattern as production.
  3. Select three workload types. Pick one stateless service, one stateful service with EBS, and one platform-heavy workload that uses observability or security agents.
  4. Run a scheduling audit. Check node selectors, affinity, topology spread, PDBs, tolerations, privileged mode, hostPath, and DaemonSets before migration.
  5. Force normal failure modes. Roll deployments, delete pods, scale replicas up and down, trigger node consolidation if possible, and simulate one Spot-tolerant workload if you plan to use Spot.
  6. Validate platform signals. Confirm logs, metrics, traces, runtime alerts, security events, load balancer provisioning, DNS, and persistent volume operations.
  7. Compare costs and toil. Record EC2 instance mix, the published Auto Mode management fee for each instance type and region, pod density, pending time, interruption behavior, and operator actions required.

Success criteria should be explicit:

  • 95%+ of pilot pods schedule without manual intervention
  • No silent loss of logs, metrics, traces, or security alerts
  • PDBs allow node replacement for replicated services
  • Stateful workloads survive rescheduling and volume attachment tests
  • Cost model is understood at instance-family level, not estimated from a generic percentage
  • Production migration blockers are documented with owners

If the pilot fails, that is still useful. It tells you which node assumptions are real and which workloads should stay on managed node groups.


FAQ

Does EKS Auto Mode work with existing EKS clusters?

Yes, Auto Mode can be enabled on existing clusters running Kubernetes 1.29 or greater, provided the cluster meets the IAM, add-on, and networking requirements. Existing managed node groups can continue to run while you migrate workloads gradually. Treat mixed operation as a transition state with clear scheduling boundaries.

Can I still use kubectl and standard Kubernetes tooling with Auto Mode?

Yes. From the workload API perspective, it is still Kubernetes. kubectl, Helm, Argo CD, Flux, policy engines, and CI/CD workflows should continue to work unless they depend on node-level implementation details.

What happens when a node AMI has a CVE?

AWS manages the node image and replacement flow for Auto Mode nodes. Your responsibility is to make sure workloads can be disrupted safely: replicas, PDBs, graceful shutdown, readiness probes, and topology rules all matter.

Can I use my existing Karpenter NodePools and EC2NodeClasses?

Not directly. Auto Mode uses Karpenter NodePool resources, but the AWS-specific node class is NodeClass under eks.amazonaws.com/v1, not the self-managed Karpenter EC2NodeClass. Review every field before porting anything.

Is EKS Auto Mode available in all AWS regions?

At launch, AWS announced Auto Mode in all AWS Regions where EKS was available except AWS GovCloud (US) and China Regions. That is no longer the full current picture: AWS later announced availability in both AWS GovCloud (US-East) and AWS GovCloud (US-West), and AWS China announced availability in the China (Beijing) and China (Ningxia) Regions. AWS documentation also lists Auto Mode AMI accounts across current commercial and GovCloud Regions. Still verify the target Region before rollout, because regional launches and partition-specific requirements can lag; AWS China, for example, documents Kubernetes 1.30 or later for Auto Mode.

Does Auto Mode support Windows nodes?

No. AWS currently states that EKS Auto Mode does not support Windows nodes. Windows workloads should stay on managed node groups or self-managed Windows nodes.

Does Auto Mode remove the need for HPA or KEDA?

No. Auto Mode handles node provisioning and lifecycle. It does not decide how many replicas your application should run. You still need HPA, KEDA, custom controllers, or application-level scaling logic for pod replica counts.

Is Auto Mode cheaper than managed node groups?

Not automatically. Auto Mode adds a management fee on top of EC2 and EKS control plane costs. It may still lower total cost if it improves bin packing, reduces over-provisioning, increases Spot usage safely, or saves meaningful platform engineering time. Measure it with your workload mix.

What is the biggest migration risk?

Hidden node assumptions. DaemonSets, privileged pods, hostPath, strict PDBs, old node labels, custom CSI drivers, and security agents are the areas most likely to break or degrade silently.


Sources