Stop Implementing Authentication Inside Containers on Kubernetes

Stop Implementing Authentication Inside Containers on Kubernetes

You have ten microservices running in Kubernetes. Each one validates JWTs, checks scopes, maintains sessions, and implements its own RBAC rules. One team uses jsonwebtoken v8, another uses a custom Go library, a third rolled their own HMAC check because “it was simple.” They all accept alg: none. Three accept RS256 and HS256 simultaneously.

This is not a security posture. This is a distributed security liability — and Kubernetes makes the problem worse, because the cluster creates an illusion of isolation that encourages teams to treat each Pod as a security boundary it was never designed to be.

The pattern of embedding authentication and authorization logic inside individual containers is one of the most pervasive anti-patterns in Kubernetes-based microservices. It feels like ownership and simplicity. It is, in practice, inconsistency at scale — and the blast radius of a single misconfiguration is your entire service portfolio.

Kubernetes provides the primitives to fix this at the infrastructure layer: Ingress controllers, the Gateway API, service mesh sidecars, admission webhooks, and workload identity via SPIFFE. None of these require a line of auth code inside your application containers.

This article explains why the anti-pattern exists, what’s wrong with it technically, and what the correct Kubernetes-native alternatives are — with concrete implementation guidance and references to the standards and incidents that validate the argument.


The Anti-Pattern: What It Looks Like

In-Application JWT Validation

Every service imports an auth library and validates tokens independently:

# Pattern seen in thousands of microservices
from jose import jwt

def authenticate(token: str):
    payload = jwt.decode(token, SECRET_KEY, algorithms=["HS256"])
    return payload["sub"]

Variations include:

  • Algorithm confusion: accepting both HS256 and RS256, or letting the token header drive verification behavior instead of pinning acceptable algorithms server-side. This is a distinct JWT implementation failure class, documented extensively by PortSwigger on JWT attacks and RFC 8725
  • alg: none bypass: libraries that accept unsigned tokens when alg is set to none. This is a documented attack vector in Auth0’s JWT security analysis
  • Missing exp / iss / aud validation: trusting a valid signature without checking whether the token is expired, for the right audience, or from the right issuer
  • Key confusion attacks: accepting a public RS256 key as an HS256 symmetric secret

Session State in Every Service

When services maintain session state directly, they duplicate logic that has no business being duplicated — cookie validation, refresh token flows, PKCE verification — and each implementation diverges over time.

RBAC Reimplemented Per Service

Authorization rules (“can this user access this resource?”) end up embedded in service logic, mixed with business logic, tested inconsistently, and impossible to audit across the portfolio.


Why This Is a Structural Problem

1. The Vulnerability Surface Scales With Your Service Count

Each new microservice is a new JWT validation surface. A single incorrect library configuration — an unvalidated alg, a missing aud check — is an authentication bypass affecting that entire service. With ten services, you have ten potential misconfigurations. With a hundred services, the probability that at least one is misconfigured approaches certainty.

The OWASP Kubernetes Security Cheat Sheet and OWASP Microservices Security Cheat Sheet both identify in-service auth as a primary attack surface in microservices environments. NIST SP 800-204 and its companion NIST SP 800-204A on DevSecOps make the same argument: security controls belong at infrastructure boundaries, not inside application code.

2. Maintenance Cost Is Multiplicative

When a JWT vulnerability is disclosed — and they are disclosed regularly — you update one library in one service. Then another. Then you discover service C is pinned to an old version because it has a transitive dependency conflict. Meanwhile the vulnerability is exploitable in production.

The CNCF Cloud Native Security Whitepaper frames this directly: security controls implemented redundantly across services create maintenance overhead that teams cannot sustain, leading to version drift and policy divergence.

3. Centralized Policy Is Impossible to Enforce

When policy is in code — even well-factored library code — it cannot be changed atomically across services. A policy update requires a coordinated deployment across every affected service. In practice, services deploy on different schedules, managed by different teams, with different testing cycles. The result is that at any given moment, some fraction of your services are running different authorization rules.

This is the core argument in Google’s BeyondCorp model and the Zero Trust Architecture guidance from NIST SP 800-207: authentication and authorization decisions should be made by a centralized, auditable policy enforcement point — not distributed across workloads.

4. Secrets Distribution Is a Problem You Don’t Need

If every service validates JWTs, every service needs the signing key (for symmetric algorithms) or the public key (for asymmetric). Distributing and rotating signing keys across a fleet of microservices is an operational burden with meaningful blast radius: a leaked symmetric key compromises every service holding it.

The CNCF SPIFFE/SPIRE project was built specifically to solve this class of problem: workload identity should be cryptographically attested, not rely on secrets distributed to application code.


The Real-World Incidents

The alg: none Class

In 2015, critical vulnerabilities in JWT libraries from Auth0 affecting Python, PHP, Node.js, Ruby, Java, and .NET allowed attackers to forge tokens by setting alg: none. The signature was not verified. The vulnerability was present in applications that had copied JWT validation code from tutorials or used unpatched libraries — exactly the pattern that in-service auth produces at scale.

Java’s Psychic Signatures (CVE-2022-21449)

CVE-2022-21449 affected ECDSA signature verification in Oracle Java SE and GraalVM, including java.security.Signature paths used by higher-level libraries. The bug allowed certain malformed ECDSA signatures to verify when they should not. JWT validation was in scope only when the deployment used an affected Java runtime and ECDSA-signed tokens, for example ES256. A gateway would help only if verification happened on a patched or unaffected runtime at the gateway instead of inside every Java service.

CVE-2023-2728 (Kubernetes Mountable Secrets Bypass)

CVE-2023-2728 was not an ImagePolicyWebhook outage behavior. It was a Kubernetes API server issue where users could use ephemeral containers to bypass the mountable secrets policy enforced by the ServiceAccount admission plugin. Clusters were affected only when the ServiceAccount admission plugin, the kubernetes.io/enforce-mountable-secrets annotation, and ephemeral containers were used together. The adjacent ImagePolicyWebhook issue was CVE-2023-2727, also involving ephemeral containers, but it is a separate CVE.

The Uber API Gateway Evolution

Uber’s engineering blog describes its API gateway as a centralized layer for routing, protocol conversion, rate limiting, load shedding, header propagation, security auditing, and user access blocking. That supports the architectural point here: high-volume platforms move cross-cutting controls into shared infrastructure. The public source does not prove that Uber migrated specifically from per-service authentication to gateway authentication, so that stronger claim should not be made.

Netflix Zuul

Netflix’s Zuul is an L7 gateway for dynamic routing, monitoring, resiliency, security, and related edge concerns. Netflix’s own Zuul posts list authentication among common edge-service uses, but they do not frame Zuul primarily as a case study in eliminating per-service auth. Treat it as evidence that authentication is a natural edge concern at scale, not as proof of a specific migration story.


The Alternatives

Use these as complementary controls, not as a single replacement for all authentication and authorization logic:

AlternativeWhen to use itWhat it solvesPrincipal trade-off
API Gateway / Edge AuthExternal API clients, public ingress, partner integrations, mixed auth mechanisms at the boundaryCentral JWT/API-key/OAuth2 validation, rate limiting, request shaping, and identity header propagation before traffic reaches servicesDoes not secure east-west service calls by itself; trusted headers require strict network boundaries
Service Mesh mTLSService-to-service traffic inside the cluster, especially across teams or sensitive domainsWorkload identity, automatic mTLS, peer authentication, and proxy-level authorization policyAdds data-plane/control-plane complexity and operational coupling to sidecars or ambient mesh components
OAuth2 ProxyBrowser-facing internal apps that need OIDC login, redirects, cookies, and session handlingDelegates login/session management to a reverse proxy and forwards authenticated identity headersBest for HTTP/browser flows; not a general machine-to-machine authorization system
OPAComplex, auditable, frequently changing authorization rulesSeparates policy decisions from application releases and can run as sidecar, service, ext_authz backend, or admission control via GatekeeperPolicy/data distribution and failure behavior must be designed deliberately
SPIFFE/SPIREMulti-cluster, multi-cloud, or meshless environments that need portable workload identityIssues short-lived workload identities without application-managed shared secretsProvides identity, not business authorization; needs registration and attestation lifecycle management

Option 1: API Gateway (Edge Auth)

An API Gateway sits at the perimeter and handles authentication before a request reaches any downstream service. Services receive pre-validated identity in a trusted header.

What it does: – Validates JWTs, API keys, OAuth2 tokens – Enforces rate limiting per identity – Strips and re-adds Authorization headers as needed – Routes to upstream services with verified identity headers

When to use it: – North-south traffic (external clients → cluster) – Mixed authentication mechanisms (JWT + API key + mTLS) at the Ingress layer – Teams that want to centralize auth policy without rolling out a full service mesh

Tools:

Gravitee.io API Gateway can be deployed on Kubernetes via its Helm chart and integrates with the Kubernetes Gateway API:

helm repo add graviteeio https://helm.gravitee.io
helm install gravitee-apim graviteeio/apim \
  --namespace gravitee \
  --create-namespace \
  --set gateway.replicaCount=2 \
  --set gateway.ingress.enabled=true \
  --set gateway.ingress.hosts[0]=api.example.com

Once deployed, authentication policies are declared on the ApiV4 CRD — no application code involved:

apiVersion: gravitee.io/v1alpha1
kind: ApiV4
metadata:
  name: payment-api
  namespace: gravitee
spec:
  name: "Payment API"
  type: PROXY
  listeners:
    - type: HTTP
      paths:
        - path: /v1/payments
      entrypoints:
        - type: http-proxy
  endpointGroups:
    - name: default
      type: http-proxy
      endpoints:
        - name: upstream
          type: http-proxy
          configuration:
            target: http://payment-service.production.svc.cluster.local:8080
  flows:
    - name: JWT validation
      enabled: true
      request:
        - policy: jwt
          enabled: true
          configuration:
            signature: RSA_RS256
            publicKeyResolver: JWKS_URL
            jwksUrl: https://idp.example.com/.well-known/jwks.json
            checkTokenRevocation: true
            requiredClaims:
              - name: aud
                value: payment-api
        - policy: rate-limit
          enabled: true
          configuration:
            rate:
              limit: 100
              periodTime: 1
              periodTimeUnit: MINUTES

Gravitee’s Kubernetes operator reconciles ApiV4 resources against the gateway, making API policy a first-class GitOps object — versioned, reviewed, and deployed the same way as any other Kubernetes manifest.

Other Kubernetes-native options: Emissary-Ingress (formerly Ambassador) with its AuthService CRD; Traefik with ForwardAuth middleware on IngressRoute resources.

Limitations: API Gateways handle north-south traffic. They don’t address east-west (service-to-service) authentication inside the cluster.


Option 2: Service Mesh (East-West mTLS + Auth)

A service mesh provides mutual TLS between every service pair and enforces authorization policy at the sidecar proxy, without any application code changes.

What it does: – Automatic mTLS between all service-to-service calls – Workload identity via X.509 certificates (SPIFFE SVIDs) – Fine-grained AuthorizationPolicy at the Envoy sidecar – JWT validation at the proxy, not the application

Istio Implementation:

Istio uses Envoy’s ext_authz filter and native RequestAuthentication + AuthorizationPolicy CRDs:

# RequestAuthentication — validate JWTs at the proxy
apiVersion: security.istio.io/v1beta1
kind: RequestAuthentication
metadata:
  name: require-jwt
  namespace: production
spec:
  selector:
    matchLabels:
      app: payment-service
  jwtRules:
    - issuer: "https://accounts.google.com"
      jwksUri: "https://www.googleapis.com/oauth2/v3/certs"
      audiences:
        - "my-api-audience"
      forwardOriginalToken: false
---
# AuthorizationPolicy — enforce after JWT validation
apiVersion: security.istio.io/v1beta1
kind: AuthorizationPolicy
metadata:
  name: payment-service-authz
  namespace: production
spec:
  selector:
    matchLabels:
      app: payment-service
  action: ALLOW
  rules:
    - from:
        - source:
            principals: ["cluster.local/ns/production/sa/order-service"]
      to:
        - operation:
            methods: ["POST"]
            paths: ["/v1/payments"]
      when:
        - key: request.auth.claims[scope]
          values: ["payments:write"]

The RequestAuthentication policy tells Envoy how to validate JWTs. The AuthorizationPolicy specifies what authenticated principals are allowed to do. Neither policy lives in application code.

The payment service receives the validated request — or a 401/403 from the proxy, before the request touches application code.

Linkerd:

Linkerd provides automatic mTLS with SPIFFE-compliant workload identity. Its policy model is simpler than Istio but sufficient for most service-to-service auth requirements:

apiVersion: policy.linkerd.io/v1beta3
kind: Server
metadata:
  name: payment-server
  namespace: production
spec:
  podSelector:
    matchLabels:
      app: payment-service
  port: 8080
---
apiVersion: policy.linkerd.io/v1beta3
kind: ServerAuthorization
metadata:
  name: order-to-payment
  namespace: production
spec:
  server:
    name: payment-server
  client:
    meshTLS:
      serviceAccounts:
        - name: order-service

This is mutual TLS + SPIFFE-based identity, enforced at the proxy. The application doesn’t implement it; the mesh does.

Istio + Envoy External Authorization:

For more complex policy (e.g., OPA integration), Envoy’s ext_authz filter delegates authorization to an external service:

apiVersion: networking.istio.io/v1alpha3
kind: EnvoyFilter
metadata:
  name: ext-authz-filter
  namespace: production
spec:
  workloadSelector:
    labels:
      app: payment-service
  configPatches:
    - applyTo: HTTP_FILTER
      match:
        context: SIDECAR_INBOUND
        listener:
          filterChain:
            filter:
              name: "envoy.filters.network.http_connection_manager"
      patch:
        operation: INSERT_BEFORE
        value:
          name: envoy.filters.http.ext_authz
          typed_config:
            "@type": type.googleapis.com/envoy.extensions.filters.http.ext_authz.v3.ExtAuthz
            grpc_service:
              envoy_grpc:
                cluster_name: outbound|9191||opa.production.svc.cluster.local
            timeout: 0.5s
            failure_mode_allow: false

The CNCF TAG Security paper on microservices security documents this architecture as the reference pattern for production Kubernetes environments.


Option 3: OAuth2 Proxy (Delegated Auth for HTTP)

OAuth2 Proxy is a reverse proxy that authenticates requests against an OAuth2/OIDC provider and passes validated identity downstream. With 14,000+ GitHub stars and active maintenance, it is the most widely deployed solution for this pattern in Kubernetes.

What it does: – Sits in front of one or more upstream services – Redirects unauthenticated requests to an OIDC provider (Keycloak, Dex, Google, GitHub, etc.) – Validates tokens, manages sessions, handles refresh – Passes X-Auth-Request-User, X-Auth-Request-Email, X-Auth-Request-Groups headers downstream

Kubernetes deployment with Nginx Ingress:

# OAuth2 Proxy deployment
apiVersion: apps/v1
kind: Deployment
metadata:
  name: oauth2-proxy
  namespace: auth
spec:
  replicas: 2
  selector:
    matchLabels:
      app: oauth2-proxy
  template:
    metadata:
      labels:
        app: oauth2-proxy
    spec:
      containers:
        - name: oauth2-proxy
          image: quay.io/oauth2-proxy/oauth2-proxy:v7.6.0
          args:
            - --provider=oidc
            - --oidc-issuer-url=https://keycloak.example.com/realms/myrealm
            - --client-id=$(CLIENT_ID)
            - --client-secret=$(CLIENT_SECRET)
            - --cookie-secret=$(COOKIE_SECRET)
            - --http-address=0.0.0.0:4180
            - --reverse-proxy=true
            - --upstream=static://202
            - --email-domain=*
            - --set-xauthrequest=true
            - --cookie-secure=true
            - --skip-provider-button=true
          env:
            - name: CLIENT_ID
              valueFrom:
                secretKeyRef:
                  name: oauth2-proxy-secrets
                  key: client-id
            - name: CLIENT_SECRET
              valueFrom:
                secretKeyRef:
                  name: oauth2-proxy-secrets
                  key: client-secret
            - name: COOKIE_SECRET
              valueFrom:
                secretKeyRef:
                  name: oauth2-proxy-secrets
                  key: cookie-secret
---
# Ingress annotation to protect a service with OAuth2 Proxy
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: protected-service
  annotations:
    nginx.ingress.kubernetes.io/auth-url: "https://oauth2-proxy.example.com/oauth2/auth"
    nginx.ingress.kubernetes.io/auth-signin: "https://oauth2-proxy.example.com/oauth2/start?rd=$escaped_request_uri"
    nginx.ingress.kubernetes.io/auth-response-headers: "X-Auth-Request-User,X-Auth-Request-Email,X-Auth-Request-Groups"
spec:
  rules:
    - host: app.example.com
      http:
        paths:
          - path: /
            pathType: Prefix
            backend:
              service:
                name: protected-service
                port:
                  number: 8080

The upstream service receives X-Auth-Request-User and X-Auth-Request-Groups as trusted headers — it never sees a token, never validates a signature, never imports a JWT library.

When to use OAuth2 Proxy vs a service mesh: OAuth2 Proxy handles north-south browser-facing traffic with session management (login flows, redirects, cookies). A service mesh handles east-west machine-to-machine auth. They are complementary, not alternatives.


Option 4: Open Policy Agent (OPA / OPAL)

OPA decouples policy from code entirely. Authorization logic is written in Rego and evaluated by OPA as a sidecar or as a centralized service. Applications query OPA for allow/deny decisions.

# Rego policy — payment service authorization
package payments.authz

import future.keywords.if
import future.keywords.in

default allow := false

allow if {
    input.method == "POST"
    input.path == "/v1/payments"
    "payments:write" in input.token.scope
    input.token.iss == "https://accounts.example.com"
}

allow if {
    input.method == "GET"
    startswith(input.path, "/v1/payments/")
    "payments:read" in input.token.scope
}

Application code becomes:

// The ONLY auth code in the application
func authMiddleware(next http.Handler) http.Handler {
    return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
        input := map[string]interface{}{
            "method": r.Method,
            "path":   r.URL.Path,
            "token":  extractToken(r),
        }
        
        result, err := opaClient.Decision(r.Context(), "payments/authz", input)
        if err != nil || !result.Allow {
            http.Error(w, "Forbidden", http.StatusForbidden)
            return
        }
        next.ServeHTTP(w, r)
    })
}

OPAL (Open Policy Administration Layer) adds real-time policy and data updates to OPA deployments — policy changes propagate to all OPA instances within seconds without redeployment.

Kubernetes deployment patterns for OPA:

As a sidecar — OPA runs in the same Pod as the application, evaluating policy over a local socket. Zero network hop, no external dependency:

# Pod template fragment
spec:
  containers:
    - name: payment-service
      image: payment-service:latest
    - name: opa
      image: openpolicyagent/opa:0.63.0
      args:
        - run
        - --server
        - --addr=localhost:8181
        - /policy
      volumeMounts:
        - name: opa-policy
          mountPath: /policy
          readOnly: true
  volumes:
    - name: opa-policy
      configMap:
        name: payment-policy

As a centralized service with Envoy ext_authz — OPA exposes a gRPC endpoint that Istio’s Envoy sidecar calls for every request. Policy is enforced at the proxy, before the application receives the request. This is the pattern used alongside Istio’s EnvoyFilter shown in the service mesh section above.

As OPA Gatekeeper — OPA Gatekeeper runs as a Kubernetes admission webhook and enforces policies at deploy time, not at runtime. It’s the right tool for preventing misconfigured workloads from being deployed — for example, rejecting any Pod spec that sets hostNetwork: true or defines auth-related environment variables directly. This is complementary to runtime auth enforcement.

OPA is used at scale by Atlassian, Goldman Sachs, Netflix, Chef, and many others, documented in OPA’s production deployments. The CNCF OPA project graduated in 2021.


Option 5: SPIFFE/SPIRE (Workload Identity)

SPIFFE (Secure Production Identity Framework for Everyone) and SPIRE solve the problem of how workloads prove their identity without distributing secrets.

SPIRE issues short-lived X.509 SVIDs (SPIFFE Verifiable Identity Documents) to workloads. Each SVID encodes a SPIFFE URI:

spiffe://example.org/ns/production/sa/payment-service

Services authenticate each other using mTLS with these certificates. No JWT library. No shared secret. No secret distribution problem. Certificate rotation happens automatically every few hours.

# SPIRE Agent DaemonSet on Kubernetes
apiVersion: apps/v1
kind: DaemonSet
metadata:
  name: spire-agent
  namespace: spire
  labels:
    app: spire-agent
spec:
  selector:
    matchLabels:
      app: spire-agent
  template:
    metadata:
      labels:
        app: spire-agent
    spec:
      serviceAccountName: spire-agent
      hostPID: true
      hostNetwork: true
      dnsPolicy: ClusterFirstWithHostNet
      containers:
        - name: spire-agent
          image: ghcr.io/spiffe/spire-agent:1.9.0
          args: ["-config", "/run/spire/config/agent.conf"]
          volumeMounts:
            - name: spire-config
              mountPath: /run/spire/config
              readOnly: true
            - name: spire-agent-socket
              mountPath: /run/spire/sockets
              readOnly: false
      volumes:
        - name: spire-config
          configMap:
            name: spire-agent
        - name: spire-agent-socket
          hostPath:
            path: /run/spire/sockets
            type: DirectoryOrCreate

SPIFFE is the foundation of Istio’s workload identity model. Linkerd implements SPIFFE-compatible SVIDs. If you’re using a service mesh, you already have SPIFFE — the mesh uses it transparently.

Standalone SPIRE is appropriate for environments without a service mesh, or for multi-cluster/multi-cloud scenarios where a consistent workload identity layer is needed across boundaries.

SPIFFE/SPIRE graduated from the CNCF sandbox to incubating in 2019 and is deployed at Uber, Bloomberg, ByteDance, and Anthem.


Combining the Layers

These tools are not mutually exclusive — they address different traffic patterns and different problems:

LayerToolAddresses
Edge (north-south)API Gateway (Gravitee, Emissary, AWS APIGW)External clients → cluster
Browser sessionsOAuth2 ProxyBrowser-facing apps with login flows
East-west mTLSService Mesh (Istio/Linkerd) + SPIFFEService-to-service identity
PolicyOPA/OPALFine-grained, auditable authorization
Workload identitySPIREMulti-cloud/multi-cluster identity

A production deployment at reasonable scale looks like:

  1. External traffic hits an API Gateway or Ingress controller with OAuth2 Proxy
  2. The gateway validates the token, strips it, and forwards identity headers to the upstream service
  3. Inside the cluster, all service-to-service calls are mTLS via a service mesh, using SPIFFE workload identity
  4. Authorization decisions (beyond identity) are delegated to OPA
  5. Application code contains zero JWT validation, zero session management, zero auth library imports

Migration Path

If you already have auth embedded in services, migration doesn’t require a big bang rewrite.

Phase 1: Introduce the Gateway

Deploy an API Gateway or OAuth2 Proxy at the edge. Initially, services continue to validate tokens themselves as a backup — the gateway validates first. Use this phase to verify the gateway’s behavior and build confidence.

Phase 2: Trust the Gateway

Add a service-level feature flag: if a trusted X-Auth-Request-User header is present (set by the gateway), skip internal JWT validation. This decouples service auth from gateway rollout.

Phase 3: Remove In-Service Auth

Once all entry points are covered by the gateway and you have confidence in its reliability, remove the auth code from services. This is the step that actually reduces your attack surface.

Phase 4: Add East-West (Optional)

If east-west service-to-service calls exist and carry sensitive data, introduce a service mesh for mTLS. This is a separate effort from gateway auth and can proceed independently.


Decision Framework

External traffic entering the cluster (Ingress / Gateway API)?
├── Browser-facing app with login flow → OAuth2 Proxy (Nginx Ingress annotations)
└── API clients with tokens → API Gateway (Gravitee ApiV4 / Emissary AuthService)

Service-to-service calls inside the cluster (east-west)?
├── mTLS + identity sufficient → Service Mesh (Istio/Linkerd)
├── Identity across multiple clusters → SPIFFE/SPIRE standalone
└── Istio + fine-grained policy → RequestAuthentication + OPA via ext_authz

Complex, auditable authorization logic?
└── OPA as sidecar or ext_authz (runtime) + OPA Gatekeeper (admission)

Preventing misconfigured workloads from being deployed?
└── OPA Gatekeeper admission webhook

What to Keep in Application Code

Not everything should be removed. The correct model is:

  • Remove: JWT signature verification, token parsing, OAuth2 flows, session management
  • Keep: Business-level authorization (“can this user edit this specific resource?”), assuming identity is provided by infrastructure
  • Keep: Authorization errors surfaced correctly (403 vs 401, meaningful error bodies)
  • Keep: Structured logging of authorization decisions for audit trails

The application receives an authenticated identity from infrastructure. What the application does with that identity — which records to show, which operations to allow based on ownership — is correctly application logic.


Audit Checklist: Moving Auth Out of Containers

Use this as a practical exit checklist for the anti-pattern:

  • Inventory every service that imports JWT, OAuth2, OIDC, session, or custom RBAC libraries.
  • Classify each entry point as north-south, browser session, east-west service call, deploy-time admission policy, or business authorization.
  • Put one enforcing control in front of each class: API Gateway or OAuth2 Proxy for ingress, service mesh mTLS for east-west, OPA for shared policy, and SPIFFE/SPIRE for portable workload identity.
  • Pin JWT issuers, audiences, algorithms, and JWKS sources in infrastructure policy; do not let application code infer them from token headers.
  • Strip client-supplied identity headers at the edge and re-add trusted identity headers only after verification.
  • Define failure behavior explicitly: fail closed for authentication and authorization, and document any temporary fail-open exception with an owner and expiry date.
  • Remove in-service token verification only after every ingress path is covered, logs prove the infrastructure control is enforcing, and rollback has been tested.
  • Keep resource-level business authorization in application code, but feed it an identity established by infrastructure.

References

Standards and Frameworks – NIST SP 800-204: Security Strategies for Microservices – NIST SP 800-204A: Building Secure Microservices-based Applications Using Service-Mesh Architecture – NIST SP 800-207: Zero Trust Architecture – OWASP Kubernetes Security Cheat Sheet – OWASP Microservices Security Cheat Sheet – CNCF Cloud Native Security Whitepaper v2 – RFC 7519: JSON Web Token (JWT) – RFC 8725: JSON Web Token Best Current Practices

Vulnerabilities and Attack Classes – CVE-2022-21449: Java Psychic Signatures (ECDSA bypass) — analysis by Neil Madden – CVE-2023-2728: Kubernetes mountable secrets policy bypass – CVE-2023-2727: Kubernetes ImagePolicyWebhook bypass – Auth0: Critical Vulnerabilities in JSON Web Token Libraries (alg:none, RS/HS confusion) – PortSwigger Web Security Academy: JWT attacks – jwt.io: Debugger and library reference

Tools and Projects – OAuth2 Proxy (GitHub — 14k+ stars) – OAuth2 Proxy Documentation – Gravitee.io API Gateway — JWT Policy – Gravitee Kubernetes Operator – Gravitee Helm Chart – Emissary-Ingress AuthService – OPA Gatekeeper – Istio Security: RequestAuthentication and AuthorizationPolicy – Envoy External Authorization Filter – Linkerd Server Policy – Open Policy Agent – OPAL — Open Policy Administration Layer – SPIFFE — Secure Production Identity Framework for Everyone – SPIRE — SPIFFE Runtime Environment – Netflix Zuul (GitHub) – Traefik ForwardAuth Middleware

Architecture and Industry Context – Google BeyondCorp: A New Approach to Enterprise Security – Google BeyondCorp Research Paper (USENIX ;login:) – Netflix Tech Blog: Zuul 2 — The Netflix Journey to Asynchronous, Non-Blocking Systems – SPIFFE/SPIRE CNCF Graduation Announcement – OPA CNCF Graduation – InfoQ: Microservices Authentication and Authorization Anti-Patterns – AWS re:Invent: Zero Trust Networking on AWS – Istio Service Mesh Security Architecture – The CNCF TAG Security Microservices Security Paper


Article reflects tooling as of 2026: Kubernetes 1.29+, Istio 1.21+, Linkerd 2.15+, OPA 0.63+ / Gatekeeper 3.16+, SPIRE 1.9+, OAuth2 Proxy 7.6+, Gravitee APIM 4.x.

Sources

Kaniko, BuildKit, and Image Volumes: The Evolution of Container Images Inside Kubernetes

Kaniko, BuildKit, and Image Volumes: The Evolution of Container Images Inside Kubernetes

Running container images inside Kubernetes is table stakes. Building them there — or mounting their contents as volumes — has been a moving target for years. What started as a privileged hack has evolved into a set of mature, secure, and increasingly native primitives.

This article covers the full arc: why the original approach was broken, what Kaniko solved, why its maintenance status now matters, where BuildKit and Buildah fit, and what Image Volumes (stable in Kubernetes v1.36) change for workflows that don’t need to build anything at all.


The original problem: Docker-in-Docker

The first generation of CI/CD on Kubernetes ran Docker inside Docker. You mounted the host Docker socket (/var/run/docker.sock) into a build container, and that container had full access to the host’s Docker daemon.

# The approach nobody should be using in 2026
volumes:
- name: docker-sock
  hostPath:
    path: /var/run/docker.sock

This worked. It was also a complete security disaster.

Mounting the Docker socket gives the container root-equivalent access to the host. Any workload that can reach that socket can escape the container, inspect other containers, and compromise the node. The only thing standing between your CI pipeline and a full cluster compromise was the good intentions of whoever wrote the build script.

It got worse when the Kubernetes project deprecated Dockershim in v1.20 and removed it in v1.24. Clusters that moved to containerd or CRI-O no longer had a Docker daemon on the host at all. The socket either didn’t exist or belonged to a completely different runtime. Docker-in-Docker in its classic form became functionally impossible on modern clusters.


Kaniko: daemonless builds inside containers

Google released Kaniko in 2018 to solve exactly this problem. Kaniko builds container images entirely in userspace — no daemon, no privileged access, no host socket required.

The key insight: Docker builds work by executing each RUN instruction in a temporary container, snapshotting the filesystem, and saving the result as a layer. Kaniko replicates this logic without a daemon. It runs as a regular container, reads the Dockerfile, executes each step against the local filesystem, and pushes the resulting image directly to a registry.

That design aged well. The project governance did not. The GoogleContainerTools/kaniko repository was archived by its owner on June 3, 2025 and is now read-only. In practical terms, the original Google-hosted Kaniko project should be treated as unmaintained: no active upstream issue triage, no normal pull request flow, and no clear path for security fixes through that repository.

apiVersion: v1
kind: Pod
metadata:
  name: kaniko-build
spec:
  containers:
  - name: kaniko
    image: gcr.io/kaniko-project/executor:latest
    args:
    - "--dockerfile=Dockerfile"
    - "--context=git://github.com/your-org/your-repo"
    - "--destination=your-registry/your-image:tag"
    volumeMounts:
    - name: registry-creds
      mountPath: /kaniko/.docker
  volumes:
  - name: registry-creds
    secret:
      secretName: registry-credentials
      items:
      - key: .dockerconfigjson
        path: config.json
  restartPolicy: Never

No host socket. No privileged flag. The container needs write access to its own filesystem (so readOnlyRootFilesystem: true won’t work), but that’s a far narrower requirement than socket mounting.

When Kaniko is still a reasonable answer

Kaniko can still be a reasonable tactical choice when you need to build a container image inside Kubernetes and already have a working, isolated pipeline around it. Existing Kaniko jobs did not stop working when the repository was archived.

For new platform work in 2026, though, do not pick Kaniko by default. The maintenance signal changes the risk model. Use it only if the simplicity is worth owning the upgrade and vulnerability-management story yourself, or if you deliberately standardize on a maintained fork or vendor-supported distribution.

Kaniko integrates well with: – Tekton — the standard pattern is a Tekton Task running the Kaniko executor – Argo Workflows — same pattern, different orchestrator – GitLab CI on Kubernetes — many older examples and pipelines use Kaniko with the Kubernetes executor – Any pod-based CI system — it’s just a container, so it runs anywhere pods run


The alternatives: BuildKit and Buildah

Kaniko is no longer the only serious daemonless option. Two other tools are worth evaluating first for new work:

BuildKit (rootless mode)

BuildKit is the build backend behind docker buildx and the default build system in modern Docker. In rootless mode it runs without privileges and can build images inside a Kubernetes pod.

BuildKit has better caching than Kaniko — particularly layer caching via cache mounts — and supports more advanced Dockerfile features like heredocs and multi-platform builds. The tradeoff is more complex setup: you need to run buildkitd as a sidecar or as a DaemonSet.

For teams already using docker buildx locally, BuildKit is the most natural migration path. Docker’s Kubernetes driver can run BuildKit builders directly in a cluster, including rootless mode without privileged pods on supported Kubernetes versions. The operational cost is real — builder lifecycle, cache persistence, node placement, and rootless kernel requirements — but the project is active and the feature set is where most modern Dockerfile workflows are moving.

Buildah

Buildah is the containers project’s daemonless build tool, designed to integrate with Podman and OCI-native workflows. It can build from Dockerfiles or Containerfiles, and it is the natural choice on OpenShift or in environments where the platform team already standardizes on Red Hat, Podman, and the containers/* stack.

The caveat is that rootless Buildah inside a restricted Kubernetes pod still depends on user namespace behavior and helper binaries such as newuidmap and newgidmap. That is manageable on platforms built for it, especially OpenShift, but it is not automatically a drop-in replacement for every generic Kubernetes CI runner.


Image Volumes: a different problem entirely

Here is where the narrative splits. Everything above is about building images. Image Volumes solve a completely different problem: consuming the contents of an OCI image as a volume, without building anything.

The feature was introduced as alpha in Kubernetes v1.31, moved to beta in v1.33 with subPath and subPathExpr support, became beta enabled by default in v1.35, and graduated to stable (GA) in Kubernetes v1.36.

What it does

Image Volumes let you reference an OCI image in a pod’s volumes section and mount its filesystem contents directly into a container:

apiVersion: v1
kind: Pod
metadata:
  name: image-volume-example
spec:
  containers:
  - name: app
    image: debian
    command: ["sleep", "infinity"]
    volumeMounts:
    - name: config-data
      mountPath: /app/config
  volumes:
  - name: config-data
    image:
      reference: your-registry/your-config-image:v1.2.0
      pullPolicy: IfNotPresent

The container at /app/config sees the contents of your-config-image:v1.2.0. The volume is read-only. No init container required.

From Kubernetes v1.33+, you can also mount a subdirectory with subPath:

volumeMounts:
- name: config-data
  mountPath: /app/config
  subPath: environments/production

The use case: OCI images as artifact bundles

This feature is motivated by a pattern that has been growing quietly: using OCI images not as runnable containers but as versioned, signed, distributable artifact bundles.

The idea: instead of storing configuration, schemas, WASM modules, ML models, or static binaries in a ConfigMap or a separate volume, you package them as an OCI image. You get:

  • Version control — image tags and digests, same tooling you already use
  • Distribution — your existing registry, your existing pull secrets, your existing access controls
  • Signing and attestation — cosign, Sigstore, the full supply chain tooling works on these artifacts
  • Immutability — a digest-pinned image reference is cryptographically immutable

Image Volumes are the Kubernetes primitive that makes this pattern first-class. Without Image Volumes, the workaround was an init container that pulled the image and copied the contents to an emptyDir. It worked, but it was boilerplate.

What it does not do

Image Volumes are not a replacement for Kaniko or any build tool. They consume images; they don’t produce them. If your workflow involves building a new image from source, you still need Kaniko, BuildKit, or equivalent.

They also require container runtime support. CRI-O supported the initial alpha implementation from its v1.31 line and tracked beta support for v1.33; containerd support landed later through the containerd 2.x line. On anything before Kubernetes v1.36, verify both the Kubernetes feature gate and the runtime version before treating this as production plumbing.


Decision framework by Kubernetes version

K8s versionBuild images (CI)Mount image contents as volume
< 1.24Avoid Docker socket builds; use BuildKit, Buildah, or legacy Kaniko with explicit risk acceptanceInit container + emptyDir workaround
1.24 – 1.30BuildKit or Buildah preferred; legacy Kaniko only if already standardizedInit container + emptyDir workaround
1.31 – 1.32BuildKit or Buildah preferred; legacy Kaniko only with maintenance planImage Volumes alpha — ImageVolume feature gate required, runtime support required
1.33 – 1.34BuildKit or Buildah preferred; legacy Kaniko only with maintenance planImage Volumes beta with subPath/subPathExpr, but disabled by default; enable ImageVolume and verify runtime support
1.35BuildKit or Buildah preferred; legacy Kaniko only with maintenance planImage Volumes beta, enabled by default; still verify runtime support before production rollout
1.36+BuildKit or Buildah preferred; legacy Kaniko only for existing pipelines or maintained forksImage Volumes stable (GA), enabled by default

When to use what

Use Kaniko if: – You already have working Kaniko pipelines and the operational cost of migration is higher than the current risk – You have a maintained fork, vendor support, or an internal patching process – You want the simplest daemonless build setup and accept that upstream Google Kaniko is archived

Use BuildKit (rootless) if: – You need advanced cache mounts or multi-platform builds – Your team already uses docker buildx locally – You’re willing to run and operate BuildKit builders in the cluster

Use Buildah if: – You are on OpenShift or a Podman/Red Hat-oriented platform – You want an OCI-native build tool that does not require a daemon – Your cluster policy supports the user namespace requirements of rootless builds

Use Image Volumes if: – You want to inject versioned, signed artifacts into pods without building anything – You’re replacing init container + emptyDir patterns – You’re adopting OCI images as a general artifact format (configs, schemas, binaries)


Conclusion

The container image story inside Kubernetes has matured significantly. Docker-in-Docker is dead — correctly so. Kaniko solved an important build problem in 2018, but the original Google project is archived in 2026, so it should no longer be the default recommendation for new platforms. BuildKit and Buildah are the healthier starting points for active build pipelines, with Kaniko reserved for existing estates or explicitly supported forks.

Image Volumes are genuinely new ground. They’re not competing with Kaniko — they’re addressing a different layer of the same ecosystem: the distribution and consumption of OCI artifacts beyond just “images you run.” With GA in v1.36 and the supply chain tooling around OCI reaching maturity, the pattern of “package it as an image, distribute it like an image, mount it where you need it” is becoming the right answer for a class of problems that previously lived in ConfigMaps, PVCs, or init container hacks.

The right combination depends on what you’re doing and what version you’re running. But for the first time, Kubernetes has native answers for both sides of the equation.

Choose today

If you operate Kubernetes v1.36 or newer, use Image Volumes for read-only artifact injection and choose BuildKit or Buildah for builds. If you are on v1.35, Image Volumes are beta and enabled by default, but you should still verify runtime support and keep the init-container pattern as the rollback path. If you are on v1.33 or v1.34, beta does not mean default-on: enable the ImageVolume feature gate deliberately and validate the runtime. If you are on v1.31 or v1.32, treat Image Volumes as alpha and non-default. If you are below v1.31, Image Volumes are not part of the platform: use emptyDir plus an init container for artifact mounting, and modernize the build path separately.

For CI builds, start with BuildKit when you want Dockerfile compatibility, cache performance, multi-platform output, and a path aligned with docker buildx. Start with Buildah when your cluster is OpenShift or Podman-oriented. Keep Kaniko only where it already works and where someone owns the maintenance risk.

Sources

  • https://github.com/GoogleContainerTools/kaniko
  • https://github.com/kubernetes/enhancements/issues/4639
  • https://kubernetes.io/docs/reference/command-line-tools-reference/feature-gates/
  • https://kubernetes.io/blog/2024/08/16/kubernetes-1-31-image-volume-source/
  • https://kubernetes.io/blog/2025/04/29/kubernetes-v1-33-image-volume-beta/
  • https://kubernetes.io/docs/tasks/configure-pod-container/image-volumes/
  • https://docs.docker.com/build/builders/drivers/kubernetes/
  • https://github.com/moby/buildkit/blob/master/docs/rootless.md
  • https://github.com/containers/buildah
  • https://github.com/containers/buildah/blob/main/docs/tutorials/05-openshift-rootless-build.md

EKS Auto Mode: Pricing, vs Managed Node Groups and Karpenter, and What Breaks

EKS Auto Mode: Pricing, vs Managed Node Groups and Karpenter, and What Breaks

EKS Auto Mode is Amazon EKS with AWS running the node layer for you: it provisions EC2 instances with Karpenter-style logic, patches and replaces them on a fixed lifetime, and runs networking, DNS, block storage and load balancing as managed capabilities instead of add-ons you install. You pay for it with a management fee per EC2 instance-hour, on top of the normal EC2 price, and you give up node-level access and custom AMIs.

That is the whole trade in one sentence. The rest of this guide is about whether it is a good trade for your cluster: how the pricing actually works (with a worked example), how Auto Mode compares to managed node groups, self-managed Karpenter and Fargate, what changed in 2025-2026, what breaks when you migrate, and a one-week pilot to decide without risking production.

What EKS Auto Mode is

EKS Auto Mode, generally available since December 1, 2024, shifts more Kubernetes infrastructure responsibility from you to AWS. You still run an EKS cluster in your AWS account, and your workloads still use the Kubernetes API, but AWS takes over much of the compute, storage, networking, node lifecycle, and core add-on management that platform teams usually wire together themselves.

For compute, Auto Mode uses Karpenter under the hood. When pods are unschedulable, Auto Mode provisions nodes that fit the workload’s requirements: instance family, size, architecture, capacity type, and availability zone. When capacity is no longer useful, it consolidates and terminates nodes. The nodes run a locked-down variant of Bottlerocket with SELinux enforcing and a read-only root filesystem, and there is no SSH or SSM access.

The important framing is this: Auto Mode is not “EKS without nodes.” It is EKS where AWS manages the node lifecycle more aggressively. You own the workloads, their scheduling requirements, their disruption behavior, and the operational consequences of those choices. AWS owns more of the infrastructure plumbing.

What it replaces

Before Auto Mode, running EKS in production usually meant choosing and operating several layers yourself:

  • Managed node groups: you chose instance types, defined scaling ranges, managed AMI updates, handled node draining, and configured Cluster Autoscaler or another scaling mechanism.
  • Self-managed Karpenter: more flexible than managed node groups, but you owned the Karpenter controller, IAM, NodePools, EC2NodeClasses, the Spot interruption queue, disruption settings, upgrades, and failure modes.
  • Fargate: AWS-managed compute per pod, with no node management, but no DaemonSets, a narrower workload compatibility envelope, and a different cost model.

EKS Auto Mode replaces a large part of that platform assembly with a managed model: declare workload intent and high-level compute constraints; AWS provisions and manages the EC2 instances behind it. If you are still deciding between the two open-source autoscalers, see Cluster Autoscaler vs Karpenter first; Auto Mode is essentially “Karpenter, operated by AWS”.

EKS Auto Mode pricing: how the fee actually works

This is the part most people search for, and the part most summaries get slightly wrong. The official structure, from the Amazon EKS pricing page (checked September 2026), is:

Monthly cost of an EKS Auto Mode cluster =
    EKS control plane            ($0.10 per cluster-hour, standard support)
  + EC2 instances                (normal price: On-Demand, Spot, RI or Savings Plans)
  + EKS Auto Mode management fee (per instance-hour, by instance type and region)
  + everything else              (EBS, load balancers, NAT, data transfer, CloudWatch...)

Auto Mode fee = Σ (Auto Mode-managed instance runtime × regional fee for that instance type)

Four details matter more than the headline:

  • The fee is per instance, not per cluster. It is charged only on EC2 instances that Auto Mode launches and manages. A cluster that runs a mix of Auto Mode nodes and legacy managed node groups only pays the fee on the Auto Mode ones.
  • It is billed like EC2: per second, with a one-minute minimum.
  • It is independent of how you buy the EC2 capacity. On-Demand, Spot, Reserved Instances and Compute Savings Plans all work with Auto Mode, but the management fee does not get discounted with them. This is the single most important pricing detail, and we will see why in the example below.
  • It varies by instance type and region. AWS does not publish it as a contractual percentage. In the public examples for US West (Oregon) the fee works out to exactly 12% of the On-Demand rate, but always pull the real per-type rate for your region before modelling anything.

The published Oregon examples are:

InstanceEC2 On-Demand ($/h)Auto Mode fee ($/h)Fee as % of On-Demand
m5a.xlarge0.1720.0206412%
m5a.2xlarge0.3440.0412812%
c6a.2xlarge0.3060.0367212%
c6a.4xlarge0.6120.0734412%

GPU and accelerated instances have their own fee schedule, and AWS cut it on July 1, 2026: management fees for G-series instances dropped by 35%, and for P-series and Trainium by 60%, automatically for every Auto Mode cluster. If your earlier cost model for an ML cluster was built in 2025, redo it.

Worked example: 10 × m5a.2xlarge for a month

Take a steady cluster of ten m5a.2xlarge nodes in us-west-2, running 730 hours a month:

Line itemManaged node groupsEKS Auto Mode
Control plane$73.00$73.00
EC2 On-Demand (10 × $0.344 × 730)$2,511.20$2,511.20
Auto Mode fee (10 × $0.04128 × 730)—$301.34
Total$2,584.20$2,885.54

On raw bill, Auto Mode is about +11.7% for this cluster. Now apply discounts to the EC2 part only, because that is what happens in reality (the discount levels below are illustrative, not quotes):

EC2 purchase optionEC2 costAuto Mode feeFee relative to your EC2 bill
On-Demand$2,511.20$301.3412%
Savings Plan at ~30% off$1,757.84$301.34~17%
Spot at ~65% off$878.92$301.34~34%

The better you already are at buying compute, the more expensive Auto Mode looks in relative terms, because the fee is anchored to the On-Demand-scale price and does not shrink with your discount. A team that runs mostly Spot pays the equivalent of roughly a third on top of its compute bill.

The other side of the ledger is efficiency. If Auto Mode’s consolidation lets the same workloads run on nine nodes instead of ten, the Auto Mode bill becomes $2,260.08 EC2 + $271.21 fee + $73 = $2,604.29, within 1% of the managed node group cluster. In other words, you need roughly 11-12% better bin packing to break even on the raw bill at On-Demand prices. Clusters with fixed-shape node groups sized “just in case” often waste more than that; clusters already on well-tuned Karpenter usually do not.

Then add the part the bill does not show. The fee on this cluster is about $3,600 a year. At 100 nodes it is about $36,000 a year. Compare that against the hours your team spends on AMI rotation, autoscaler upgrades, add-on version matrices and node incidents. For a two-person platform team the fee is usually cheap; for a large platform team that already automated all of it, it is usually not.

Fargate for comparison

Fargate is priced per pod, per second, on requested vCPU and memory. In US East (N. Virginia), Linux/x86 is about $0.04048 per vCPU-hour and $0.004446 per GB-hour (September 2026), so a 1 vCPU / 2 GB pod running all month costs about $36. There is no node to pay for, but also nothing to bin-pack: you pay for requests, not usage, which makes right-sized requests and limits even more important than on EC2. And Amazon EKS does not support Fargate Spot.

EKS Auto Mode vs managed node groups vs Karpenter vs Fargate

EKS Auto ModeManaged node groupsSelf-managed KarpenterFargate
Who runs node provisioningAWS (managed Karpenter)You (ASG + Cluster Autoscaler)You (Karpenter controller)AWS, per pod
Instance selectionDynamic via NodePoolsFixed list per node groupDynamic via NodePoolsNot exposed
Node OS / AMIBottlerocket variant, chosen by AWSAL2023, Bottlerocket, Windows or customAL2023, Bottlerocket, Windows or customNot exposed
AMI patchingAutomatic, weekly AMIs, 21-day max node lifetimeYou roll updatesYou set expireAfter and driftN/A
SSH / SSM to nodesNoYes, if enabledYes, if enabledNo
DaemonSetsYes, but no host changesYesYesNo
WindowsNoYesYesNo
SpotYes, interruptions handled nativelyYesYes, you configure the SQS interruption queueNo
Networking, DNS, EBS CSI, LB controllerManaged capabilitiesAdd-ons you install and upgradeAdd-ons you install and upgradeManaged, with Fargate limits
Extra costManagement fee per instance-hourNoneNone (controller runs on your capacity)Per-pod vCPU/GB pricing
Operational burdenLow for nodes, medium for workload compatibilityMediumMedium-highLow for nodes, medium for compatibility
Right forTeams that don’t need node customizationRegulated/custom node environmentsMature platform teams that want full controlIsolated, simple workloads

EKS Auto Mode vs managed node groups

This is the comparison most teams actually face. Managed node groups give you a fixed capacity shape and full control of the node: your AMI, your kubelet flags, your host agents. Auto Mode gives you dynamic capacity and removes node operations, but takes the node away from you.

Pick managed node groups if you need custom or hardened AMIs, Windows, host-level security tooling, specific kernel settings, or if your capacity is steady and already heavily discounted (the fee hurts most there). Pick Auto Mode if your node groups are over-provisioned, your team spends real time on AMI rotations and add-on upgrades, and your workloads tolerate nodes being replaced at least every 21 days.

The two can coexist in the same cluster, which is also how migrations work: new workloads on Auto Mode, legacy or special ones on node groups.

EKS Auto Mode vs Karpenter

Auto Mode is Karpenter, but not your Karpenter. Both share the karpenter.sh/v1 NodePool API; the node class differs (NodeClass in eks.amazonaws.com/v1 for Auto Mode, EC2NodeClass in karpenter.k8s.aws/v1 for self-managed), and so do the instance labels (eks.amazonaws.com/instance-family vs karpenter.k8s.aws/instance-family).

Self-managed Karpenter has no fee and lets you choose the AMI, OS (including Windows), and node lifetime. In exchange you operate the controller, its IAM, the Spot interruption queue, GPU drivers and device plugins, and every upgrade. Auto Mode handles Spot interruptions without an SQS queue, ships NVIDIA and Neuron drivers, enables node repair by default, and upgrades the controller for you. If you already run Karpenter well, the fee buys you little. If you were about to adopt Karpenter, Auto Mode removes most of the adoption cost.

How it works in practice

You create or update an EKS cluster with Auto Mode enabled. The default setup uses two AWS-managed built-in node pools, system and general-purpose, which you can enable or disable but not modify. If you need more control, you create a NodeClass for Auto Mode infrastructure settings and a NodePool for workload-facing scheduling constraints.

# NodeClass: EKS Auto Mode infrastructure settings for managed EC2 nodes.
apiVersion: eks.amazonaws.com/v1
kind: NodeClass
metadata:
  name: private-compute
spec:
  subnetSelectorTerms:
    - tags:
        kubernetes.io/role/internal-elb: "1"
  securityGroupSelectorTerms:
    - tags:
        aws:eks:cluster-name: prod-eks
  ephemeralStorage:
    size: "100Gi"
# NodePool: workload-facing constraints for nodes that Auto Mode may provision.
apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
  name: general-purpose-custom
spec:
  template:
    metadata:
      labels:
        billing-team: platform
    spec:
      nodeClassRef:
        group: eks.amazonaws.com
        kind: NodeClass
        name: private-compute
      requirements:
        - key: karpenter.sh/capacity-type
          operator: In
          values: ["on-demand", "spot"]
        - key: kubernetes.io/arch
          operator: In
          values: ["amd64", "arm64"]
        - key: eks.amazonaws.com/instance-category
          operator: In
          values: ["c", "m", "r"]
  limits:
    cpu: "1000"
    memory: 1000Gi
  disruption:
    consolidationPolicy: Balanced
    consolidateAfter: 1m

The NodeClass is where you express AWS infrastructure placement and node-level defaults. The NodePool is where you express what kind of capacity is acceptable for workloads. Do not copy a self-managed Karpenter EC2NodeClass into Auto Mode; Auto Mode uses its own NodeClass API.

A few defaults are worth knowing because they drive both disruption and cost. Auto Mode consolidates underutilized nodes, expires nodes after 336 hours (14 days) unless you change it, never lets a node live beyond 21 days, applies a disruption budget of 10% of nodes, and replaces nodes for drift when a new Auto Mode AMI ships, roughly once a week. consolidationPolicy accepts WhenEmpty, WhenEmptyOrUnderutilized and Balanced; the last one weighs disruption cost against savings and skips consolidations that are not worth it, which is a sensible default for most production pools. A billing label in the NodePool template (like billing-team above) is the cheapest way to get per-team cost allocation later.

Built-in components are managed differently than in a classic EKS build. Pod networking (with network policy enforcement), service networking, cluster DNS, compute autoscaling, EBS block storage, load balancing, the Pod Identity agent and node monitoring are Auto Mode capabilities. With Auto Mode compute, add-ons such as Amazon VPC CNI, kube-proxy, CoreDNS, the Amazon EBS CSI Driver and the EKS Pod Identity Agent become redundant for Auto Mode nodes, and the controllers run on AWS-owned infrastructure rather than as visible pods in your account. Autoscaling of replicas is still yours: Auto Mode adds nodes for pending pods, but you still need an HPA or KEDA to create those pods (see HPA best practices).

The workload support matrix is broader than Fargate, but not identical to self-managed nodes:

CapabilityAuto Mode status
EC2 SpotSupported through karpenter.sh/capacity-type (spot, on-demand, reserved)
Graviton / arm64Supported through kubernetes.io/arch: arm64
GPU / acceleratorsSupported for documented NVIDIA, Trainium and Inferentia families; drivers and device plugins managed
Capacity reservationsODCRs and Capacity Blocks for ML via capacityReservationSelectorTerms
Windows nodesNot supported (Linux only)
DaemonSetsSupported, but host-level assumptions must be validated against locked-down nodes

What changed in 2025 and 2026

Auto Mode has moved fast since launch, and several of the early “reasons not to use it” are gone. The ones worth knowing, all from AWS announcements:

  • October 2025, networking and security controls. The NodeClass gained podSubnetSelectorTerms and podSecurityGroupSelectorTerms for separate pod subnets, advancedNetworking options for public IP control and forward HTTPS proxies (httpsProxy, noProxy), certificateBundles for enterprise CAs, KMS encryption of root and ephemeral volumes, and support for On-Demand Capacity Reservations and Capacity Blocks for ML. Corporate proxies and private CAs used to be hard blockers; they are not anymore.
  • October 2025, GovCloud. Auto Mode became available in AWS GovCloud (US-East) and (US-West), and later in the AWS China Regions.
  • February 2026, enhanced logging. Auto Mode’s managed capabilities (compute autoscaling, block storage, load balancing, pod networking) can be configured as CloudWatch Vended Logs sources and delivered to CloudWatch Logs, S3 or Firehose. This closes the “black box” gap: you can finally see why the managed Karpenter did or did not launch a node.
  • April 2026, hidden managed instances. New Auto Mode instances and their resources are hidden by default from EC2 console views and API list operations. Good for accidental-deletion safety, surprising for cost and inventory tooling that lists EC2 instances directly. Adjust the managed resource visibility settings if your FinOps scripts rely on them.
  • July 2026, cheaper accelerated compute. Management fees cut by 35% for G-series and 60% for P-series and Trainium, effective July 1.
  • July 2026, zonal shift. Integration with Application Recovery Controller zonal shift: during a zonal shift Auto Mode stops provisioning in the impaired AZ and pauses voluntary disruption, at no extra cost.
  • July 2026, EFA and placement groups for Auto Mode node pools, aimed at distributed training.
  • Static capacity NodePools. Setting replicas on a NodePool keeps a fixed number of nodes regardless of demand, useful for latency-sensitive inference. Once set it cannot be removed, and static pools are never consolidated, so set limits.nodes above replicas to leave headroom for node replacement.

What you gain

Reduced operational surface. Node group management, AMI lifecycle, Cluster Autoscaler tuning, Karpenter controller upgrades, and a chunk of add-on wiring move out of your day-to-day scope.

Better provisioning shape by default. Dynamic provisioning is usually a better fit than fixed node group shapes. You get nodes that more closely match actual pod requirements instead of trying to pre-plan a small set of instance types, and that is exactly the efficiency that has to pay for the fee.

Automatic node patching. AWS manages the node image and replacement flow. That reduces toil, but it also means your workloads need disruption policies that let AWS replace nodes safely.

Faster cluster bootstrapping. A new Auto Mode cluster gets to a usable production baseline faster than a hand-assembled EKS cluster with node groups, autoscaling, networking add-ons, storage drivers, and load balancer controllers.

Native Spot integration. Spot interruptions, rebalance recommendations, scheduled maintenance and instance health events are handled for you, with no Node Termination Handler or SQS queue. You still need workload-level interruption tolerance: replicas, budgets, graceful shutdown, and queue semantics where relevant.

What you give up

Node-level access. Auto Mode nodes are intentionally locked down. If your incident response process assumes SSH, SSM, manual package inspection, or ad hoc host changes, it needs to change. Ephemeral debug containers become your main tool; the patterns in debugging distroless containers apply almost verbatim.

Custom AMIs. AWS determines the operating system and AMI for Auto Mode managed instances; you cannot directly access the instance or install software on it. If your organization requires internally built, hardened, or certified AMIs, Auto Mode is likely blocked. (For how Bottlerocket compares to other container-optimized OSes, see Talos vs Bottlerocket vs Flatcar.)

Unrestricted host agents. Kubernetes DaemonSets are supported, but they are the sharp edge. Anything that assumes privileged host access, custom kernel modules, hostPath writes, or low-level runtime integration needs a proof of compatibility.

Less tuning surface. You give up direct control over kubelet flags, container runtime configuration, bootstrap scripts, and arbitrary node setup. That is the point of the product, but it is also the boundary.

Different cost visibility. Managed node groups make capacity easy to reason about because you chose it up front. Auto Mode changes capacity dynamically and, since April 2026, hides its instances from default EC2 listings, so cost control moves toward NodePool limits, billing labels, Cost Explorer and workload-level resource hygiene.

When NOT to use EKS Auto Mode

Auto Mode is a good default, but there are clear cases where it is the wrong answer:

  • You need Windows nodes. Auto Mode is Linux only.
  • You must own the node image supply chain: certified or hardened AMIs, CIS images built in-house, or specific kernel modules.
  • Your security or observability stack needs host access that the locked-down Bottlerocket nodes do not allow, and the vendor has no supported Auto Mode deployment.
  • Your compute is already heavily discounted and efficiently packed. With mostly Spot or Savings Plans and a well-tuned Karpenter, the fee is a large relative premium for work you have already automated.
  • You run singleton stateful workloads that cannot tolerate node replacement at least every 21 days. Auto Mode will rotate the node anyway; if the PDB blocks it, you are just postponing an incident.
  • You rely on an alternative CNI or networking setup that Auto Mode does not support.
  • You need tight kubelet or runtime tuning (custom eviction thresholds, runtime configuration, special sysctls on the host).

When to use EKS Auto Mode

Use Auto Mode if:

  • You run EKS on AWS and do not have a hard requirement to manage nodes yourself
  • You do not require custom AMIs or custom node bootstrap logic
  • You want to reduce the operations surface for node lifecycle management
  • You want Karpenter-like provisioning without operating Karpenter yourself
  • Your workloads are mostly stateless or disruption-tolerant
  • Your observability, security, and storage agents are compatible with Auto Mode

Stick with managed node groups if:

  • Your organization requires internally certified or hardened AMIs
  • You need specific kernel configuration, kubelet flags, bootstrap scripts, or host packages
  • You depend on privileged DaemonSets or host-level security tooling that Auto Mode cannot support
  • You are in a regulated environment where the node image supply chain must be owned internally
  • Your platform team already operates Karpenter well and values the extra control

Use Fargate if:

  • You specifically want per-pod compute isolation
  • Your workload does not need DaemonSets or host-level integrations
  • You accept Fargate’s scheduling, networking, storage, and observability constraints
  • You want to avoid managing EC2 capacity entirely for a narrow class of workloads

Migration from managed node groups

Migrating an existing cluster to Auto Mode is supported, but it is not a one-command operational migration. You must update the cluster IAM role permissions and trust policy, enable the compute, block storage, and load balancing capabilities together, and meet required add-on versions when those add-ons are installed. AWS does not support direct migration of EBS volumes from the standard EBS CSI provisioner to the Auto Mode one, of existing load balancers from AWS Load Balancer Controller to Auto Mode, or of clusters using alternative CNIs. A conservative path looks like this:

  1. Enable Auto Mode on a non-production cluster running a supported Kubernetes version (1.29 or later).
  2. Inventory workloads by scheduling assumptions: node selectors, affinities, tolerations, topology spread, PDBs, privileged mode, hostPath, local storage, and DaemonSet dependencies.
  3. Create or select the relevant Auto Mode NodeClass and NodePool resources.
  4. Move a low-risk namespace first by changing selectors, tolerations, or labels so pods land on Auto Mode nodes (eks.amazonaws.com/compute-type: auto identifies them).
  5. Watch scheduling, replacement, load balancer behavior, persistent volume provisioning, logging, metrics, and security events.
  6. Taint old node groups to stop new scheduling once the pilot workloads are stable.
  7. Drain old nodes gradually and delete old managed node groups only after workload owners have signed off.

You can keep AWS Load Balancer Controller installed during the migration when you need both models; separate them with IngressClass or loadBalancerClass and plan a blue-green move for existing load balancers.

The hardest part is usually not enabling Auto Mode. It is discovering which workloads and platform agents quietly depended on a mutable node.

What breaks when you migrate

Auto Mode changes the node contract. The Kubernetes API still looks familiar, but the host underneath is no longer yours in the same way.

DaemonSets need a compatibility audit. Logging agents, metrics agents, service mesh node components, security scanners, CSI node plugins, and custom infrastructure daemons often assume host access. Datadog, Falco, custom CSI drivers, eBPF agents, file integrity tools, and in-house node agents should be tested explicitly rather than assumed compatible.

PodDisruptionBudgets can block AWS-managed maintenance. If every critical Deployment has maxUnavailable: 0, or singleton workloads have no safe disruption path, node replacement becomes harder, and the 21-day lifetime means it will happen anyway. Auto Mode can manage nodes, but it cannot make an application disruption-tolerant after the fact.

nodeSelector and affinity rules can strand pods. Workloads pinned to old node group labels, instance types, capacity labels, zones, or custom AMI labels may never schedule on Auto Mode capacity. Replace legacy labels with stable requirements Auto Mode can satisfy, and remember that instance labels use the eks.amazonaws.com/ prefix, not karpenter.k8s.aws/.

topologySpreadConstraints can become too strict. Auto Mode provisions capacity dynamically, but strict zone spreading plus narrow selectors can create unschedulable pods. Check whenUnsatisfiable, label selectors, and minimum domain assumptions.

Privileged pods and hostPath volumes are migration blockers until proven otherwise. Anything that needs /var/lib, /proc, /sys, container runtime sockets, kernel capabilities, or host networking deserves a separate test. For a broader list of patterns that hurt under node churn, see Kubernetes features that hurt in production.

Observability and security agents may lose host assumptions. The dangerous failure mode is not CrashLoopBackOff; it is silent loss of telemetry or enforcement.

Storage drivers must be reviewed. EBS integration is part of Auto Mode, but it uses the provisioner ebs.csi.eks.amazonaws.com, not the standard ebs.csi.aws.com. Custom CSI drivers, EFS patterns, snapshot controllers, and topology-aware storage classes should be validated. Pay attention to provisioner names, volume binding mode, encryption settings, and IAM assumptions.

Runbooks need rewriting. “SSH to the node and inspect X” is not a valid first response anymore. Incident procedures should move toward kubectl describe, events, logs, ephemeral debug containers, the Auto Mode vended logs, cloud-side metrics, and vendor-supported diagnostics.

Cost and inventory scripts may go blind. Since April 2026, Auto Mode instances are hidden from default EC2 list operations. Anything that inventories or tags instances by listing EC2 directly needs the visibility setting changed or a switch to Kubernetes-side data.

1-week pilot: evaluate Auto Mode without risking production

Use a short pilot to answer compatibility and economics questions before touching production.

Day 1: build a representative test cluster. Same region, Kubernetes minor version, VPC shape, IAM model, ingress pattern, and storage classes where possible.

Day 1-2: enable Auto Mode with one custom NodeClass and NodePool. Keep the first pool boring: on-demand capacity, two or three common instance families, the same private subnet pattern as production, and a billing label.

Day 2-3: move three workload types. One stateless service, one stateful service with EBS, and one platform-heavy workload that uses observability or security agents. Run the scheduling audit (selectors, affinity, topology spread, PDBs, tolerations, privileged mode, hostPath, DaemonSets) before each move.

Day 4: force normal failure modes. Roll deployments, delete pods, scale replicas up and down, trigger consolidation, and simulate one Spot-tolerant workload if you plan to use Spot.

Day 5: validate platform signals. Confirm logs, metrics, traces, runtime alerts, security events, load balancer provisioning, DNS, and persistent volume operations. Enable the Auto Mode vended logs and check that you can explain every node launch.

Day 6-7: compare costs and toil. Record the EC2 instance mix, the published fee for each instance type in your region, pod density, pending time, interruption behavior, and operator actions required. Rerun the worked example above with your real numbers and your real discounts.

Success criteria should be explicit:

  • 95%+ of pilot pods schedule without manual intervention
  • No silent loss of logs, metrics, traces, or security alerts
  • PDBs allow node replacement for replicated services
  • Stateful workloads survive rescheduling and volume attachment tests
  • Cost model is understood at instance-type level, including the fee on discounted capacity
  • Production migration blockers are documented with owners

If the pilot fails, that is still useful. It tells you which node assumptions are real and which workloads should stay on managed node groups.

The honest assessment

EKS Auto Mode is a good default candidate for many teams running Kubernetes on AWS, and a better one in late 2026 than at launch: proxies, private CAs, capacity reservations, logging and zonal shift were real gaps that have been closed. The operational simplification is real: node provisioning, AMI updates, core add-on integration, and scaling behavior are areas where teams burn time and create incidents.

The constraints are also real. Custom AMIs, Windows, unrestricted host access from DaemonSets, privileged pods, custom CSI drivers, and strict disruption policies are the common blockers. And the fee is not trivial for teams that already buy compute cheaply: it does not shrink with Spot or Savings Plans.

The right question is not whether Auto Mode is “better” than managed node groups. It is which operational contract your workloads can live with, and whether the fee is smaller than the node-management work it removes. The more your organization can treat nodes as replaceable capacity, the more value you get. The more your platform treats nodes as customized machines, the less Auto Mode fits.

Frequently Asked Questions

How much does EKS Auto Mode cost?

You pay the normal EKS control plane fee ($0.10 per cluster-hour on standard support), the normal EC2 price for the instances, and an EKS Auto Mode management fee per instance-hour that depends on instance type and region, billed per second with a one-minute minimum. In AWS’s published US West (Oregon) examples the fee equals 12% of the On-Demand price (for example $0.04128/hour on an m5a.2xlarge), but always check the rate for your instance types and region on the EKS pricing page.

Do Savings Plans, Reserved Instances or Spot reduce the EKS Auto Mode fee?

No. You can use On-Demand, Spot, Reserved Instances and Compute Savings Plans for the EC2 capacity that Auto Mode launches, and those discounts apply to the EC2 part of the bill, but the Auto Mode management fee is independent of the purchase option. That is why the fee is a larger relative premium for teams that already run mostly Spot or discounted capacity.

Is EKS Auto Mode cheaper than managed node groups?

Not on the raw bill for the same capacity: managed node groups have no management fee. Auto Mode can still lower total cost if its consolidation and right-sized instance selection remove enough idle capacity (around 11-12% better packing breaks even at On-Demand prices) or if it saves meaningful platform engineering time. Measure it with a pilot and your real discounts.

What is the difference between EKS Auto Mode and Karpenter?

Auto Mode runs Karpenter for you. Both use the same karpenter.sh/v1 NodePool API, but Auto Mode uses its own NodeClass (eks.amazonaws.com/v1) instead of EC2NodeClass, a locked-down Bottlerocket AMI chosen by AWS, native Spot interruption handling and managed GPU drivers. Self-managed Karpenter has no fee and lets you choose AMIs, OS and node lifetime, but you operate the controller, IAM, interruption queue and upgrades.

Can I enable EKS Auto Mode on an existing cluster?

Yes. Auto Mode can be enabled on existing clusters on a supported Kubernetes version, provided the cluster meets the IAM, add-on and networking requirements. Existing managed node groups keep running while you move workloads gradually. EBS volumes from the standard CSI provisioner and existing AWS Load Balancer Controller load balancers do not migrate directly and need a planned cut-over.

Can I use my existing Karpenter NodePools and EC2NodeClasses with Auto Mode?

Not directly. NodePools are the same API and usually port with small changes (instance labels use the eks.amazonaws.com/ prefix), but the self-managed EC2NodeClass has no equivalent in Auto Mode; you need to recreate its settings as an Auto Mode NodeClass and review every field before porting anything.

Does EKS Auto Mode support Windows nodes or SSH access?

No to both. Auto Mode supports only Linux nodes, and the nodes do not allow SSH or SSM access. Windows workloads should stay on managed node groups or self-managed nodes, and troubleshooting moves to Kubernetes-level tools such as ephemeral debug containers, events, logs and the Auto Mode vended logs.

Does EKS Auto Mode replace HPA or KEDA?

No. Auto Mode handles node provisioning and lifecycle: it adds nodes when pods are pending and removes them when they are not needed. It does not decide how many replicas your application should run. You still need HPA, KEDA or application-level scaling for pod replica counts.

Sources

Skopeo, Crane, and regctl: Container Image Management Without the Docker Daemon (2026)

Skopeo, Crane, and regctl: Container Image Management Without the Docker Daemon (2026)

The Problem: Docker Is Overkill for Image Operations

You need to copy an image from Docker Hub to your private registry. Or inspect a manifest before pulling. Or delete old tags programmatically. Or sync an entire repository during a migration.

Related reading: Kaniko, BuildKit and Image Volumes.

The instinct is to reach for Docker. But Docker requires a running daemon, root access (or group membership that amounts to the same thing), and pulls the entire image to disk just to read its metadata. For CI pipelines, GitOps workflows, and platform tooling, that’s a significant overhead for what should be lightweight registry operations.

This is the problem that daemonless container image tools solve. Skopeo pioneered the category; today it has real competition from crane, regctl, and ORAS — each with different strengths and ideal use cases.

This article gives you the practical comparison to pick the right tool for your workflow.


The Contenders

ToolMaintainerLanguageDaemon required
SkopeoRed Hat / containersGoNo
craneGoogle / ko-buildGoNo
regctlregclientGoNo
ORASCNCFGoNo
cosignSigstore / OpenSSFGoNo

All five are Go binaries, statically compiled, and work directly against the OCI Distribution Spec. None of them require Docker or any container runtime.


Skopeo

Skopeo was the first major tool to address daemonless image operations, released by Red Hat in 2016 as part of the containers/image ecosystem (alongside Podman and Buildah).

What Skopeo does well

Image inspection without pulling:

skopeo inspect docker://registry.k8s.io/pause:3.9

Returns full image metadata — digest, layers, labels, architecture, OS — without downloading a single layer. Useful in admission controllers, policy checks, and pre-deployment validation.

Cross-registry copying:

skopeo copy 
  docker://docker.io/library/nginx:1.27 
  docker://harbor.internal/library/nginx:1.27

Copies image manifests and layers directly between registries, bypassing your local machine entirely. The image never touches your disk.

Multi-arch handling:

skopeo copy --all 
  docker://docker.io/library/nginx:1.27 
  docker://harbor.internal/library/nginx:1.27

The --all flag copies the full manifest list, preserving all architectures (linux/amd64, linux/arm64, etc.). This is critical when mirroring images for multi-arch clusters.

Registry synchronization:

skopeo sync 
  --src docker 
  --dest docker 
  --all 
  docker.io/library/nginx 
  harbor.internal/mirrors/

skopeo sync mirrors an entire repository, including all tags. You can also use a YAML file to define which images and tags to sync — useful for air-gapped environment bootstrapping.

Tag deletion:

skopeo delete docker://harbor.internal/myapp:old-tag

Useful in CI for cleanup pipelines. Note that registry-side deletion requires the registry to have the DELETE method enabled.

Skopeo’s weaknesses

  • No image modification: Skopeo copies and inspects, but doesn’t build or modify images
  • Tag listing is verbose: skopeo list-tags returns JSON you need to parse
  • No retry logic by default: transient network errors in long sync operations require wrapping with retry scripts
  • Auth configuration: relies on containers/auth.json format, which differs from Docker’s ~/.docker/config.json (though it supports both)

When to use Skopeo

  • Air-gapped environment image mirroring
  • CI pipelines that need to copy or inspect images without Docker
  • Platform teams on Red Hat / OpenShift stacks
  • Any workflow already using Podman or Buildah

Crane

Crane is Google’s answer to Skopeo, developed as part of the ko project and later extracted into its own tool. It’s simpler, more scriptable, and has a cleaner CLI design.

What Crane does well

Tag listing:

crane ls registry.k8s.io/pause

No JSON parsing needed. One tag per line. Pipe directly into grep, sort, head.

Digest resolution:

crane digest docker.io/library/nginx:1.27

Returns the image digest. Combine with yq or sed to pin image references in Helm values or Kubernetes manifests.

Image copying:

crane cp docker.io/library/nginx:1.27 harbor.internal/library/nginx:1.27

Same capability as Skopeo’s copy, arguably with a cleaner syntax.

Manifest inspection:

crane manifest docker.io/library/nginx:1.27 | jq .

Returns raw manifest JSON. Useful when you need the exact manifest for digest verification or policy enforcement.

Tagging and retagging:

crane tag harbor.internal/myapp:abc123 harbor.internal/myapp:stable

Adds a new tag to an existing image without re-uploading layers. The tag operation is purely a manifest pointer update.

Flattening images:

crane flatten docker.io/library/ubuntu:24.04 -t harbor.internal/ubuntu:flat

Squashes all layers into one. Reduces layer count for images where layer history doesn’t matter.

Crane’s weaknesses

  • No sync command: unlike Skopeo, crane has no built-in repository sync. You script it yourself with crane ls + crane cp in a loop
  • Less mature multi-arch support: crane cp supports multi-arch but the UX is less explicit than Skopeo’s --all
  • No delete command: doesn’t implement registry deletion

When to use Crane

  • CI/CD scripting where you want clean, pipeable output
  • Digest pinning workflows
  • Lightweight image tagging operations
  • When you’re already in the ko / Google Cloud ecosystem

regctl

Regctl is the least known of the three but arguably the most feature-complete. It’s the CLI for the regclient Go library and covers use cases that Skopeo and crane leave out.

What regctl does uniquely well

Image modification without rebuild:

regctl image mod myimage:tag 
  --label "org.opencontainers.image.version=1.2.3" 
  --replace

You can add/change labels, annotations, and config fields directly on an existing image in the registry — without pulling, rebuilding, or pushing a new image. This is impossible with Skopeo or crane.

Layer operations:

# Remove a specific layer from an image
regctl image mod myimage:tag 
  --layer-rm sha256:abc123... 
  --replace

Useful for removing accidentally included secrets or large unnecessary layers from published images.

OCI artifact support:

regctl artifact put 
  --media-type application/vnd.example.config.v1+json 
  --config config.json 
  file.tar.gz 
  harbor.internal/myartifacts:v1

Regctl has solid OCI artifact support alongside standard image operations.

Formatting and output:

regctl tag list harbor.internal/myapp --format '{{range .}}{{println .}}{{end}}'

Go template formatting throughout. Useful for integrating into shell scripts without jq.

Referrers (OCI 1.1):

regctl manifest get-list harbor.internal/myapp:v1 --referrers

Lists referrers (signatures, SBOMs, attestations) attached to an image via the OCI 1.1 referrers API.

Regctl’s weaknesses

  • Smaller community: fewer examples, less StackOverflow coverage
  • Steeper learning curve: more commands, more flags
  • Less packaging: not in most distro repos by default

When to use regctl

  • Image post-processing (labels, annotations, layer removal) without rebuild
  • Advanced manifest and referrer workflows
  • When you need OCI artifact operations alongside image operations

ORAS

ORAS (OCI Registry As Storage) is a CNCF project focused specifically on OCI artifact management — pushing and pulling arbitrary files to container registries, not necessarily container images.

# Push a Helm chart as an OCI artifact
oras push harbor.internal/charts/myapp:1.0.0 
  --artifact-type application/vnd.helm.chart.v1+tar 
  mychart.tgz

# Push SBOM
oras push harbor.internal/myapp:v1 
  --artifact-type application/spdx+json 
  sbom.spdx.json

# Pull
oras pull harbor.internal/charts/myapp:1.0.0

ORAS is not a direct Skopeo replacement — it’s for when your registry is a general-purpose artifact store, not just a container registry. Helm OCI, SBOMs, attestations, and policy bundles all benefit from ORAS.

ORAS vs Skopeo in one line: Skopeo (and crane) move container images between registries; ORAS pushes and pulls arbitrary artifacts (charts, SBOMs, ML models) to a registry. If your object is an OCI image, use Skopeo or crane. If it is a file you want to store in a registry, use ORAS. They overlap far less than the shared “registry client” label suggests, and most platform teams end up with both installed for different jobs.


cosign

Cosign from Sigstore is not a general-purpose image tool — it’s specifically for supply chain security. But it’s increasingly part of any container image workflow.

# Sign an image
cosign sign --key cosign.key harbor.internal/myapp:v1@sha256:abc123...

# Verify
cosign verify --key cosign.pub harbor.internal/myapp:v1

# Attach SBOM
cosign attach sbom --sbom sbom.spdx harbor.internal/myapp:v1

# Keyless signing (Sigstore)
cosign sign harbor.internal/myapp:v1

Cosign integrates with OIDC providers for keyless signing (no key management required), which is the direction the ecosystem is moving. If you’re building a supply chain security practice, cosign is mandatory, not optional.


Side-by-side comparison

OperationSkopeoCraneregctl
Inspect imageinspectmanifestmanifest get
Copy imagecopycpimage copy
Copy all archescopy --allcp (auto)image copy
Sync repositorysyncscript itscript it
List tagslist-tags (JSON)ls (plain)tag list
Delete tag/imagedelete—tag delete
Modify labels——image mod
Remove layer——image mod
OCI artifactslimitedlimitedartifact
Referrers (1.1)——manifest get-list

Practical workflows

Copy an image from Docker Hub to a private registry (crane cp vs skopeo copy)

This is the single most common reason people reach for these tools, so here are the two one-liners side by side. Both copy the image directly registry-to-registry — the image never touches your local disk and no Docker daemon is involved.

# skopeo copy — explicit docker:// transport on both sides
skopeo copy
  docker://docker.io/library/nginx:1.27
  docker://harbor.internal/library/nginx:1.27

# crane cp — shorter, no transport prefix
crane cp docker.io/library/nginx:1.27 harbor.internal/library/nginx:1.27

They are functionally equivalent for a single image. Reach for skopeo copy when you want --all to force the full multi-arch manifest list, or when you are on a Red Hat / Podman stack. Reach for crane cp when you want the shortest possible command and cleaner output for scripting — it copies the manifest list automatically when the source is multi-arch. To authenticate against the private registry first, both read ~/.docker/config.json, so a prior docker login harbor.internal (or crane auth login / skopeo login) is enough. To copy many images or whole repositories in one go, jump to the air-gapped mirroring recipe below, which uses skopeo sync.

Mirror images for air-gapped clusters (Skopeo)

# sync-list.yaml
docker.io:
  images:
    library/nginx:
      - "1.25"
      - "1.26"
      - "1.27"
    library/redis:
      - "7.2"
      - "7.4"
skopeo sync 
  --src yaml 
  --dest docker 
  --all 
  sync-list.yaml 
  harbor.internal/mirrors/

Pin image digests in CI (Crane)

#!/bin/bash
# Update image digests in values.yaml
for image in nginx:1.27 redis:7.4; do
  digest=$(crane digest docker.io/library/${image})
  echo "docker.io/library/${image}@${digest}"
done

Combine with yq to update Helm values files automatically, ensuring reproducible deployments.

Retag without re-pushing (Crane or regctl)

# After a successful deploy to staging, promote to production
crane tag harbor.internal/myapp:${GIT_SHA} harbor.internal/myapp:production

No layer transfer. The operation is a metadata update in the registry.

Add OCI annotations post-build (regctl)

regctl image mod harbor.internal/myapp:v1.2.3 
  --annotation "org.opencontainers.image.source=https://github.com/org/repo" 
  --annotation "org.opencontainers.image.revision=${GIT_SHA}" 
  --replace

Attaches build metadata to an image already in the registry, without a rebuild.

Supply chain security pipeline

# 1. Build and push
docker buildx build --push -t harbor.internal/myapp:${GIT_SHA} .

# 2. Generate SBOM
syft harbor.internal/myapp:${GIT_SHA} -o spdx-json > sbom.spdx.json

# 3. Attach SBOM
cosign attach sbom --sbom sbom.spdx.json harbor.internal/myapp:${GIT_SHA}

# 4. Sign (keyless with OIDC in CI)
cosign sign harbor.internal/myapp:${GIT_SHA}

# 5. Verify in admission controller or deployment pipeline
cosign verify 
  --certificate-identity-regexp="https://github.com/org/repo" 
  --certificate-oidc-issuer="https://token.actions.githubusercontent.com" 
  harbor.internal/myapp:${GIT_SHA}

Installation

All tools install as single static binaries:

# Skopeo (via package manager)
brew install skopeo                    # macOS
dnf install skopeo                     # RHEL/Fedora
apt install skopeo                     # Debian/Ubuntu

# Crane
brew install crane
# or binary release
curl -sL https://github.com/google/go-containerregistry/releases/latest/download/go-containerregistry_Linux_x86_64.tar.gz | tar xz crane

# regctl
curl -sL https://github.com/regclient/regclient/releases/latest/download/regctl.linux.amd64 -o /usr/local/bin/regctl
chmod +x /usr/local/bin/regctl

# ORAS
brew install oras

# cosign
brew install cosign

Which tool should you use?

Use Skopeo if: you’re on a Red Hat / OpenShift stack, you need repository sync, or you’re building air-gapped environment pipelines. It’s the most battle-tested and widely packaged.

Use Crane if: you’re scripting image operations in CI and want clean, composable CLI output. crane ls + crane cp + crane digest cover 80% of automation use cases with minimal friction.

Use regctl if: you need to modify images post-build, work with OCI referrers, or want the most complete feature set for a registry client. It has a higher learning curve but can replace both Skopeo and crane for advanced workflows.

Use ORAS if: you’re using a registry to store non-image artifacts — Helm charts, SBOMs, policy bundles, ML models.

Use cosign regardless of which of the above you pick, as soon as supply chain security matters to your organization. It’s not a replacement for the others — it’s a complement.

In practice, most platform teams end up using 2-3 of these tools together. Crane for day-to-day scripting, Skopeo for sync jobs, cosign for signing, ORAS for artifact storage.


FAQ

Can I use these tools with private registries?

Yes. All support standard registry authentication. Crane and Skopeo both read from ~/.docker/config.json. Regctl has its own config file (~/.regctl/config.json) but can import Docker credentials. Set DOCKER_CONFIG to point to your credentials file in CI environments.

Do these work with Docker Hub rate limits?

Yes, and they’re often more efficient than the Docker CLI because they only fetch manifest metadata for inspect operations, not full layers. For heavy pull workloads, authenticate with your Docker Hub credentials to get higher rate limits.

What about ECR, GCR, and Azure Container Registry?

All tools support these with the appropriate credential helpers. For ECR, use docker-credential-ecr-login. Crane has native ECR support via the --platform flag and crane auth commands. Skopeo supports ECR via --creds "AWS:$(aws ecr get-login-password)".

Are these tools safe to run in Kubernetes pods?

Yes. Since they require no daemon and no elevated privileges for read operations, they’re well-suited to run as init containers or sidecar containers in Kubernetes. Skopeo is commonly used in image pre-pulling init containers. Use a dedicated service account with least-privilege registry credentials.

Can I copy a multi-arch image and keep all platforms?

Skopeo: skopeo copy --all. Crane: crane cp copies the index automatically when the source is a manifest list. Regctl: regctl image copy preserves manifest lists by default.

Radar: A New Kubernetes IDE Worth Knowing About (vs OpenLens, FreeLens)

Radar sweep showing OpenLens, FreeLens and Radar as a Kubernetes IDE comparison

Related reading: FreeLens vs OpenLens vs Lens: which one to standardize on.

If you’ve been following Kubernetes tooling, you’ve probably already been through the Lens saga: Lens went commercial, OpenLens emerged as the community fork, then FreeLens appeared when OpenLens maintenance slowed. The pattern is familiar — a useful desktop tool, a licensing decision, a fork, another fork. Radar is not a fork. It’s a different approach to the same problem: giving engineers a useful interface for Kubernetes clusters without the friction of kubectl for every task. Built by Skyhook (YC-backed, Google Cloud Partner), it’s been live since 2025, has 1.7k+ GitHub stars, releases weekly, and the founder reaches out to the community directly. That’s usually a good signal that someone is genuinely building in public. This article covers what Radar actually does, where it pulls ahead of OpenLens and FreeLens, and when those tools are still the right choice.

The State of Kubernetes Desktop Tooling in 2026

Before getting into Radar specifically, it’s worth naming the landscape clearly:
  • Lens — the original. Electron-based, polished, now commercial (Mirantis). The free Personal tier is non-commercial only. Pro is ~$22-35/user/month.
  • OpenLens — the community fork of Lens before Mirantis closed exec/logs/shell in v6.3 (January 2023). Maintenance has slowed significantly. No active release cadence.
  • FreeLens — a more active community fork, filling the gap left by OpenLens’ decline. Restores the missing features. No commercial backing.
  • k9s — terminal TUI, fast, keyboard-driven, single-cluster. Different audience.
  • Headlamp — CNCF Sandbox project, plugin-extensible, web-based.
  • Radar — Go binary, Apache 2.0, team-oriented, topology and event timeline focused.
The problem with OpenLens and FreeLens is not that they’re bad tools — they’re genuinely useful for the solo developer with one or two clusters. The problem is that they’re single-cluster-at-a-time desktop apps with no concept of team, no persistent state, and no awareness of the modern Kubernetes ecosystem (ArgoCD, Flux, Karpenter, KEDA). As your infrastructure grows, you outgrow them.

What Radar Actually Is

Radar is available in two forms:
  • Radar OSS — a single ~30MB Go binary, Apache 2.0, free forever. Can run locally (desktop app) or deployed in-cluster via Helm. No sidecars, no feature gates.
  • Radar Cloud — same binary, adds a hosted control plane with fleet aggregation, 30-day event retention, SSO/SCIM, scoped RBAC, and shared URLs for team incident response. Priced per cluster ($99/cluster/month for Team), not per user.
The per-cluster pricing is a deliberate design decision — teams don’t pay more as they add engineers, only as they add clusters. For a 20-person platform engineering team managing 5 clusters, Radar Cloud runs $495/month. The equivalent Lens Pro seats would cost $2,200-4,200/month. For most self-hosted environments, the OSS version is sufficient and costs nothing.

Key Features

Topology View

This is the most visually distinctive feature. Radar renders a live service graph for your cluster: deployments, services, ingresses, cross-namespace dependencies, and east-west traffic flows — all in a single view without running kubectl get all -A and stitching the output together mentally. OpenLens and FreeLens have resource list views. They show you what exists. Radar shows you how things connect — which is what you actually need when debugging why Service A can’t reach Service B.

Persistent Event Timeline

Kubernetes events are ephemeral by default — they expire after approximately one hour. When something breaks at 2am and you’re looking at it at 9am, the events that explain what happened are gone. Logs may still be there if you’re running a log aggregator, but the Kubernetes-level events (pod restarts, scheduling failures, node pressure events, probe failures) are gone. Radar retains events. The OSS version extends this beyond the default 1-hour cluster retention. The Cloud version retains 30 days. You can rewind the timeline to any point and reconstruct what the cluster looked like at that moment. Neither OpenLens nor FreeLens have any event retention beyond what the cluster itself provides.

GitOps Integration (ArgoCD + Flux)

Radar auto-detects ArgoCD and Flux and surfaces sync state, drift, and health directly in the UI. You can see whether a deployment is in sync, when it last synced, and whether it drifted from the desired state in Git. In OpenLens and FreeLens, ArgoCD resources appear as generic Kubernetes custom resources. You can see the CRDs, but there’s no purpose-built understanding of what they mean — no sync status visualization, no diff view, no rollback trigger.

Helm Management

Radar tracks Helm releases with full revision history and supports one-click rollbacks from the UI. This is similar to what OpenLens/FreeLens offer via the Helm releases view, but Radar adds revision diffing — you can see what changed between release 5 and release 6 before deciding to roll back.

Image Filesystem

You can browse container image filesystems through Radar without needing kubectl exec into a running pod or access to the container registry. Useful for security audits and debugging — you can verify what’s actually in an image at rest.

MCP Server (AI Integration)

Radar ships with an MCP (Model Context Protocol) server, which means you can connect Claude, Cursor, or GitHub Copilot directly to your cluster context and ask questions about it in natural language. The MCP server is token-optimized — it doesn’t dump raw YAML at the model, it structures cluster state into meaningful context. This is something neither OpenLens nor FreeLens have. It’s also something that’s genuinely useful if you’re already using AI assistants for development work.

Cluster Audit

30 built-in best-practice checks — resource requests/limits, RBAC permissions, image pinning, network policies, security contexts. The checks are labeled by compliance framework. This is not a replacement for dedicated security tooling (Trivy, Falco, Polaris), but it’s a useful first-pass audit without leaving the tool you’re already using.

Multi-Cluster Support (Cloud)

The Cloud tier adds fleet-level visibility: a single view across all clusters, cross-cluster search, and drift detection between environments (e.g., staging vs. production). This is the feature that changes the calculus for platform engineering teams managing 5+ clusters. OpenLens and FreeLens require you to switch cluster context manually. There is no fleet view.

Architecture: Why a Go Binary Matters

OpenLens and FreeLens are Electron apps — Chromium + Node.js wrapped in a desktop shell. This means:
  • 200-500MB install size
  • 1-2 second startup time on a fast machine, more on slower ones
  • Memory footprint in the hundreds of megabytes
  • Local kubeconfig required on each engineer’s machine
Radar’s in-cluster deployment is a single Go binary (~30MB) that runs as a Pod with a ServiceAccount. It connects to the hosted control plane over outbound WebSocket + TLS. No inbound firewall rules, no kubeconfig distribution, no per-engineer setup. The local desktop app is also a lightweight Go binary — 65-second startup was demonstrated on a 322-node cluster. That’s not a typo. For in-cluster deployment, the architecture means security is handled at the ServiceAccount level, not by distributing kubeconfigs to engineer laptops. That matters for teams with security requirements around credential management.

Feature Comparison

FeatureRadar OSSRadar CloudOpenLensFreeLens
LicenseApache 2.0Proprietary (hosted)MIT/GPLMIT
MaintenanceActive (weekly releases)ActiveStalledActive (community)
ArchitectureGo binary / in-clusterIn-cluster + hostedElectronElectron
Multi-clusterBasicFleet view❌❌
Event retentionExtended30 daysCluster default (~1h)Cluster default (~1h)
Topology view✅✅❌❌
GitOps (ArgoCD/Flux)✅✅CRDs onlyCRDs only
Helm management✅✅✅✅
kubectl exec / logs / shell✅✅✅ (restored)✅
MCP / AI integration✅✅❌❌
Cluster audit✅✅❌❌
SSO / SCIM❌✅❌❌
Shared incident URLs❌✅❌❌
Image filesystem browser✅✅❌❌
Cost tracking✅ (OpenCost)✅❌❌
PriceFree$99/cluster/monthFreeFree

When Radar Makes Sense

You’re managing multiple clusters. Even with the OSS version, the topology view and event timeline make Radar more useful than OpenLens/FreeLens at 3+ clusters. The Cloud fleet view is the compelling option at 5+. Your team uses GitOps. If ArgoCD or Flux is part of your workflow, Radar’s native understanding of sync state and drift is meaningfully better than seeing CRDs in a generic list view. You need post-mortem capability. If your incident review process involves looking at what the cluster was doing when the alert fired, you need event retention. Radar has it; OpenLens and FreeLens don’t. You’re adopting AI tooling. The MCP server is the most forward-looking feature here. If you use Claude Code, Cursor, or Copilot for your infrastructure work, having cluster context available to those tools without copy-pasting YAML is a genuine productivity improvement. You have a platform engineering team. Per-cluster pricing, SSO, SCIM, and shared incident URLs are features that only matter if you have more than one person managing infrastructure.

When OpenLens or FreeLens Still Makes Sense

You’re a solo developer with one or two clusters. OpenLens and FreeLens are familiar, local, and have zero setup overhead. If you don’t need team features, event retention, or topology views, they remain perfectly functional tools. You’re deeply invested in the Lens UX. The resource tree, the terminal integration, the way Lens presents namespace-scoped resources — if your muscle memory is built around that interface, switching has a real cost. Radar is different, not just better. You need maximum customization. OpenLens and FreeLens support plugins. Radar does not currently have a plugin system. Your environment is air-gapped or has strict egress restrictions. Radar OSS can run fully in-cluster, but Radar Cloud requires outbound connectivity to the hosted control plane. OpenLens and FreeLens are fully local.

Getting Started

OSS installation takes about two minutes:
# Homebrew (macOS/Linux)
brew install skyhook-io/tap/radar

# Helm (in-cluster)
helm repo add skyhook https://charts.skyhook.io
helm install radar skyhook/radar 
  --namespace radar 
  --create-namespace 
  --set service.type=ClusterIP
Or download the binary directly from radarhq.io.

Verdict

Radar is the most interesting new entrant in the Kubernetes tooling space in a while — not because it replaces everything else, but because it addresses the specific gap that OpenLens and FreeLens never covered: teams, multiple clusters, and persistent state. For a solo developer, OpenLens or FreeLens are still completely reasonable choices. For a platform engineering team managing more than two clusters with ArgoCD or Flux, Radar’s feature set is materially better and the OSS version costs nothing. The active release cadence and the YC backing suggest this isn’t a one-person side project — there’s a team actively working on it. Whether the Cloud pricing sticks long-term is a question only usage will answer, but the Apache 2.0 core with an explicit “always open source” commitment is the right foundation. Worth evaluating if you haven’t already.
Tested with Radar OSS v0.x on Kubernetes 1.29–1.32. Pricing and feature availability as of May 2026.

Related reading: FreeLens extensions: the complete catalogue and how to install them.

Kubernetes Cluster Autoscaler vs Karpenter: When to Use Each (2026)

Kubernetes Cluster Autoscaler vs Karpenter: When to Use Each (2026)
Your pods are pending. Your on-call engineer is getting paged. Somewhere in the chain between “I need more compute” and “compute is available,” something is too slow. That something is almost always node provisioning — and the tool you chose to manage it determines whether that delay is 4 minutes or 45 seconds. Node autoscaling is one of those infrastructure decisions that looks simple until you’re running it in production. Two schedulable pods sitting in Pending state doesn’t just mean a delayed deployment — it means latency spikes, dropped traffic, breached SLOs, and engineers debugging things that should have been invisible. At scale, it also means either burning money on over-provisioned nodes or gambling on under-provisioning at the worst possible moment. Cluster Autoscaler (CA) has been the default answer for years. Karpenter emerged from AWS in 2021, graduated to stable in 2023, and by 2025 had become the default recommendation for most AWS-native clusters. In 2026, both tools are mature, widely deployed, and genuinely good — but they solve the problem differently, and picking the wrong one for your environment has real consequences. This article is a deep technical comparison. It assumes you already know what Kubernetes is and have opinions about infrastructure. The goal is to give you a clear picture of how each tool works, where each one wins, and a decision framework you can actually use.

Karpenter vs Cluster Autoscaler: the short answer

Cluster Autoscaler scales node groups you defined in advance; Karpenter provisions individual nodes on demand, choosing the instance type itself. That single architectural difference drives everything else. Cluster Autoscaler adds nodes to an existing ASG or MIG, so provisioning takes roughly 4–8 minutes and your instance choice is fixed by whatever you put in the node group. Karpenter talks to the cloud provider API directly, bin-packs pending pods against the whole instance catalogue, and typically has capacity in 60–90 seconds — while consolidating underused nodes to cut cost. Use Cluster Autoscaler if you need mature multi-cloud support, you have strict node-group governance, or your capacity is reserved and homogeneous. Use Karpenter if you are on AWS (or now Azure), your workloads are heterogeneous, and you want faster scale-up plus automatic cost consolidation. Whether you phrase it “Cluster Autoscaler vs Karpenter” or the other way round, the decision comes down to three questions — cloud, workload diversity, and how much control you want over instance selection — and there is a full decision framework at the end of this article.

Why Node Autoscaling Is Hard

The fundamental tension in autoscaling is this: you want compute available before you need it, but you don’t want to pay for compute you’re not using. These goals are in direct conflict, and every autoscaling system is an attempt to find the least-bad trade-off. Without autoscaling, you’re doing one of two things:
  1. Over-provisioning — you run enough nodes to handle peak load at all times. Your average utilization sits at 20–30%, and you’re paying for the other 70–80% to sit idle.
  2. Under-provisioning — you run lean, and when traffic spikes, pods go Pending. Your SLOs breach. You get paged at 3am to manually scale.
A common failure mode with poorly tuned autoscaling is the “thundering herd at scale-up” pattern: HPA creates new pods faster than node autoscaling can provision capacity. The provisioning window matters. With CA and typical ASG-backed node groups on AWS, you’re looking at 4–8 minutes. With Karpenter, 60–90 seconds. At 100 RPS and a 3-minute window, that’s 18,000 requests under degraded conditions.

Cluster Autoscaler: How It Actually Works

Cluster Autoscaler is a Kubernetes-native project under the kubernetes/autoscaler repository, in production since 2016, supporting AWS, GCP, Azure, Alibaba, DigitalOcean, and more.

The Node Group Model

CA operates on node groups — ASGs on AWS, MIGs on GCP, VMSSs on Azure. CA’s job is to decide when to increase or decrease the desired capacity of these groups. CA does not provision individual nodes. It scales node groups, and the node group provisions nodes. This indirection adds latency and reduces flexibility.

Scale-Up: Detecting Unschedulable Pods

CA runs a control loop (default scan interval: 10 seconds). For each Pending pod with PodScheduled=False, CA simulates adding a node of each known node group type and checks if the pod would become schedulable. When a node group is selected, CA applies an expander to choose which group to scale:
  • least-waste — minimizes CPU/memory waste after scheduling (best default for cost)
  • most-pods — maximizes pods scheduled per scale-up operation
  • priority — lets you define ordering via ConfigMap
  • grpc — delegates to an external gRPC service
# Cluster Autoscaler deployment — AWS, production-tuned
apiVersion: apps/v1
kind: Deployment
metadata:
  name: cluster-autoscaler
  namespace: kube-system
spec:
  replicas: 1
  selector:
    matchLabels:
      app: cluster-autoscaler
  template:
    metadata:
      labels:
        app: cluster-autoscaler
      annotations:
        cluster-autoscaler.kubernetes.io/safe-to-evict: "false"
    spec:
      priorityClassName: system-cluster-critical
      serviceAccountName: cluster-autoscaler
      containers:
        - image: registry.k8s.io/autoscaling/cluster-autoscaler:v1.36.1
          name: cluster-autoscaler
          resources:
            requests:
              cpu: 100m
              memory: 600Mi
            limits:
              cpu: 200m
              memory: 1Gi
          command:
            - ./cluster-autoscaler
            - --cloud-provider=aws
            - --node-group-auto-discovery=asg:tag=k8s.io/cluster-autoscaler/enabled,k8s.io/cluster-autoscaler/my-cluster
            - --expander=least-waste
            - --balance-similar-node-groups=true
            - --scale-down-delay-after-add=10m
            - --scale-down-unneeded-time=10m
            - --scale-down-utilization-threshold=0.5
            - --max-graceful-termination-sec=600
            - --scan-interval=10s

Scale-Down: The Conservative Approach

A node is a scale-down candidate only if: – CPU and memory utilization (by requests) is below threshold (default: 50%) – All pods could be rescheduled elsewhere – No pod has cluster-autoscaler.kubernetes.io/safe-to-evict: "false" – The node has been underutilized for at least --scale-down-unneeded-time (default: 10m) This conservatism prevents churn — a feature, not a bug.

Karpenter: How It Actually Works

Karpenter was built by AWS and donated to the Kubernetes project in 2023, where it lives as a SIG Autoscaling subproject under kubernetes-sigs (it has no CNCF maturity level of its own); it reached GA (v1.0) in mid-2024. Providers exist for AWS (stable), Azure (stable), and GCP (beta).

The Core Insight: Bypass the Node Group

Karpenter calls the EC2 RunInstances API directly — no ASG involvement. This means: – Any instance type in a single request, without pre-configuring a node group – No intermediary: Karpenter → EC2 API → node joins cluster – Right-size nodes to exactly what workloads need, across the full instance catalog – Karpenter handles full node lifecycle, including termination

NodePool and EC2NodeClass

# EC2NodeClass — cloud-specific parameters
apiVersion: karpenter.k8s.aws/v1
kind: EC2NodeClass
metadata:
  name: default
spec:
  amiSelectorTerms:
    - alias: al2023@latest
  role: "KarpenterNodeRole-my-cluster"
  subnetSelectorTerms:
    - tags:
        karpenter.sh/discovery: "my-cluster"
  securityGroupSelectorTerms:
    - tags:
        karpenter.sh/discovery: "my-cluster"
  blockDeviceMappings:
    - deviceName: /dev/xvda
      ebs:
        volumeSize: 50Gi
        volumeType: gp3
        encrypted: true
  metadataOptions:
    httpTokens: required
    httpPutResponseHopLimit: 1
---
# NodePool — intent and constraints
apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
  name: general-purpose
spec:
  template:
    spec:
      nodeClassRef:
        group: karpenter.k8s.aws
        kind: EC2NodeClass
        name: default
      requirements:
        - key: karpenter.sh/capacity-type
          operator: In
          values: ["on-demand", "spot"]
        - key: kubernetes.io/arch
          operator: In
          values: ["amd64", "arm64"]
        - key: karpenter.k8s.aws/instance-category
          operator: In
          values: ["c", "m", "r"]
        - key: karpenter.k8s.aws/instance-generation
          operator: Gt
          values: ["5"]
        - key: karpenter.k8s.aws/instance-size
          operator: NotIn
          values: ["nano", "micro", "small", "medium", "large"]
      expireAfter: 720h
  limits:
    cpu: "1000"
    memory: 1000Gi
  disruption:
    consolidationPolicy: WhenEmptyOrUnderutilized
    consolidateAfter: 5m
    budgets:
      - nodes: "5%"
        schedule: "0 8 * * mon-fri"
        duration: 10h
      - nodes: "25%"

Just-in-Time Provisioning and Bin Packing

When pods go Pending, Karpenter watches the event (not polls) and immediately: 1. Collects all Pending pods 2. Simulates bin packing — fewest possible nodes across the full instance catalog 3. Selects instances that satisfy all pod requirements 4. Calls EC2 API to provision the optimal instance(s)

Disruption and Consolidation

Karpenter’s differentiated value: active cluster consolidation. It evaluates whether nodes can be removed by redistributing pods onto others, or replaced with a smaller instance type. A c5.4xlarge running 4 vCPU worth of pods gets replaced with a c5.xlarge. Teams commonly report 30–50% compute cost reduction. consolidationPolicy options: – WhenEmpty — only remove nodes with no workload pods (safest) – WhenEmptyOrUnderutilized — also replace underutilized nodes with smaller ones

Architecture Comparison

DimensionCluster AutoscalerKarpenter
Node provisioning modelScales node groups (ASG/MIG/VMSS)Direct cloud API, no node groups
Instance flexibilityPre-defined node group typesFull instance catalog at runtime
Scale-up triggerPolling (10s scan interval)Watch-based event (near-instant)
Scale-downRemoves underutilized nodesRemoves + consolidates + right-sizes
Spot handlingVia ASG + AWS Node Termination HandlerNative, first-class, no NTH needed
Configuration modelDeployment flagsDeclarative CRDs
Cloud supportAll major + on-premAWS (GA), Azure (stable), GCP (beta)
ConsolidationNoYes
Community maturityVery mature (since 2016)Mature (GA 2024)

Scaling Speed: The Numbers

Cluster Autoscaler on AWS (typical): 1. Pod Pending → CA scan detects (0–10s) 2. ASG UpdateAutoScalingGroup API call (~15–30s) 3. EC2 instance starts (1–2 min) 4. Node bootstrap + kubelet registration (30–60s) 5. Pod scheduled (5–10s) Total: 3–6 minutes (up to 12 min during high-demand periods) Karpenter on AWS (typical): 1. Pod Pending → watch event fires (~1s) 2. EC2 RunInstances API call (~2–3s) 3. EC2 instance starts (same hardware — 1–2 min) 4. Node bootstraps + Ready (30–60s) 5. Pod scheduled (~5s) Total: 90 seconds to 3 minutes

Cost Optimization: Where Karpenter Pulls Ahead

Right-sizing: CA requires pre-defined node groups. Karpenter selects the minimum viable instance for the pending workload from the full catalog. Consolidation vs scale-down: CA removes underutilized nodes. Karpenter replaces a large underutilized node with a smaller one that still fits all pods. This produces compounding savings over time. Spot handling: Karpenter receives EC2 interruption notices, pre-provisions a replacement, and drains the node — all within the 2-minute window. No AWS Node Termination Handler required. It also diversifies spot requests across instance types automatically to reduce simultaneous interruption risk.

What’s New in Karpenter v1.14 (July 2026)

If your mental model of Karpenter was formed around v1.0, three changes shipped in v1.14.0 (11 July 2026) are worth revisiting, because two of them close gaps that used to be genuine reasons to stay on Cluster Autoscaler. 1. Dynamic Resource Allocation (DRA) support. Karpenter now ships a DRA allocator: it understands ResourceClaim objects when deciding what to provision, including pod-level claims, and supports consumable capacity and partitionable devices. In practice this is the GPU story. Under the old model, “this pod needs a slice of an A100” was expressed through labels, taints and affinity rules, and the autoscaler could not reason about it properly — you pinned GPU workloads to a hand-built node group and accepted the waste. With DRA, the device requirement is a first-class scheduling input, so Karpenter can pick the right accelerator instance and pack partitionable devices instead of stranding a whole GPU per pod. If you run ML workloads, this is the release that matters. 2. The CapacityBuffer API (v1beta1). Karpenter’s great weakness was always the cold start: it is fast, but it still provisions after pods go Pending. CapacityBuffer lets you declare headroom that Karpenter keeps warm, which is the “overprovisioning with pause pods” hack that teams have been hand-rolling for years, now a supported API with its own status metrics. It directly narrows the gap against a deliberately over-provisioned Cluster Autoscaler node group. 3. The Balanced consolidation policy. Consolidation used to be a fairly blunt trade — pack tighter, accept more disruption. The new Balanced policy sits between aggressive consolidation and leaving nodes alone, and consolidateAfter now also applies to destination nodes during consolidation, which makes churn considerably more predictable. If you evaluated Karpenter, liked the cost numbers and rejected it because the disruption was too noisy for your workloads, that objection is worth re-testing.

Where that leaves the release line today: v1.14.1 (21 August 2026) is the current release and is tagged as an LTS line by the AWS provider, with support committed through July 2027; the previous LTS is v1.9.x. There is no v1.15 yet. If you are choosing a version to standardise on for the next year, v1.14 is the one — it has the DRA allocator, CapacityBuffer and the Balanced policy, and it will keep receiving patches after the interim minors are dropped.

Related reading: what CNCF maturity levels actually mean.

Related reading: Kubernetes HPA memory examples.


What’s New in Cluster Autoscaler (2026)

It would be easy to read the section above as “Karpenter moves, Cluster Autoscaler stands still”. That is not what the release history says. CA tracks Kubernetes minors — 1.35.0 shipped in February 2026 and 1.36.0 in July 2026, with 1.36.1 the current patch — and two of the 2026 additions are direct answers to Karpenter’s headline features. 1. DRA support, including partitionable devices. Cluster Autoscaler can now simulate scheduling for pods that use ResourceClaims rather than the classic nvidia.com/gpu extended resource, and 1.36 added support for partitionable devices (the MIG-style “slice of an accelerator” case). It is still behind a feature flag and it still works at node-group granularity — CA can only add a node of a type you have already defined — but the “CA cannot reason about DRA at all” objection is gone. If you run GPU workloads on a non-AWS cloud, this matters more than anything Karpenter shipped. 2. CapacityBuffer v1beta1 — the same idea, on the CA side. 1.35 released a CapacityBuffer API for declaring warm headroom, now namespaced and integrated with ResourceQuotas so buffers respect the quotas you already have. This is the same problem Karpenter’s CapacityBuffer solves, arriving in the same year; the pause-pod overprovisioning hack is now a supported API on both sides of this comparison. 3. CapacityQuota and “salvo” scale-up (1.36). CapacityQuota is a CRD that caps the resources CA is allowed to scale up — a guard rail that used to require external tooling. The experimental --salvo-scale-up flag lets CA perform several scale-ups in a single loop instead of one per iteration (budgeted with --salvo-scale-up-budget), which directly attacks the “4-8 minutes” number in the scaling-speed section when many node groups need to grow at once. 4. Operational tuning. A separate --max-node-startup-time, --predicate-parallelism to run scheduler predicates on more threads (replacing --cluster-snapshot-parallelism), a “suspended” node state in the status ConfigMap, CSI volume-limit awareness during scale-up, and atomic size increases for Azure VMSS so a partially-fulfilled scale-up no longer leaves you with the wrong node count. --scale-down-enabled is deprecated in favour of the newer scale-down flags — check your Helm values before upgrading past 1.35. None of this changes the architecture: CA still scales node groups, not nodes. But if your reason for wanting Karpenter was DRA, warm capacity or scale-up throughput rather than instance flexibility, re-read the CA release notes before you migrate.

Multi-Cloud Support in 2026

CloudCluster AutoscalerKarpenter
AWS✅ Production-stable✅ Production-stable (reference impl.)
GCP✅ Production-stable⚠️ Beta (karpenter-provider-gcp)
Azure✅ Production-stable✅ Stable (karpenter-provider-azure)
Alibaba✅ Supported❌ No provider
DigitalOcean✅ Supported❌ No provider
On-premises / Cluster API✅ Supported❌ Not supported

Karpenter vs HPA, VPA and KEDA: Which Autoscaler Does What

A recurring confusion in the comparison searches: Karpenter is not an alternative to HPA or KEDA. They operate on different layers and are designed to run together. Pod autoscalers decide how many replicas you need; node autoscalers decide what hardware those replicas land on. Karpenter and Cluster Autoscaler only ever see the outcome of the pod autoscalers — pending pods.

So “KEDA vs Karpenter” and “HPA vs Karpenter” are not real choices. KEDA and HPA answer how many pods; Karpenter and Cluster Autoscaler answer how many nodes, and which ones. You will almost always run one from each row. The real comparison is HPA vs KEDA on the pod layer (KEDA when the signal is a queue, a topic or a schedule rather than CPU) and Karpenter vs Cluster Autoscaler on the node layer, which is what the rest of this article is about.

ComponentLayerWhat it changesInput signalScales to zero
HPAPodNumber of replicasCPU, memory, custom/external metricsNo (min 1)
KEDAPodNumber of replicas (drives an HPA underneath)Event sources: queue depth, Kafka lag, cron, 60+ scalersYes
VPAPodRequests and limits of each podHistorical usageN/A
Cluster AutoscalerNodeSize of predefined node groupsUnschedulable podsNode groups to 0
KarpenterNodeIndividual nodes and their instance typeUnschedulable pods + consolidationYes

The chain in practice

A queue backs up, KEDA sees the lag and raises the replica count, the new pods cannot be scheduled because the cluster is full, and Karpenter provisions a node that fits them. Every layer does its own job. The failure mode people hit is tuning only one of them: an HPA that scales aggressively on a cluster with a slow node autoscaler produces pods that sit pending for minutes, and the latency you were trying to fix stays exactly where it was.

Two pairings deserve care. HPA and VPA on the same metric fight each other — VPA raises requests, which lowers measured CPU utilisation, which makes HPA scale down; run VPA in Off or Initial mode if an HPA already targets CPU on that workload. And KEDA scaling to zero only works if the node autoscaler can drain the last node: with Karpenter, that is consolidation doing its job; with Cluster Autoscaler, you need the node group’s minimum set to 0. See HPA best practices for the pod-level half of this.

Replacing Cluster Autoscaler with Karpenter

If the goal is a straight replacement rather than coexistence, the order matters: install Karpenter with a NodePool that is restricted to the workloads you are moving, scale the Cluster Autoscaler node groups down as Karpenter takes over, and only then remove the Cluster Autoscaler deployment. Removing it first leaves the old node groups with nobody managing scale-down, and you pay for idle nodes until someone notices.


When Cluster Autoscaler Is Still the Right Choice

  1. Non-AWS environments — GCP, Alibaba, DigitalOcean, on-prem with Cluster API
  2. Existing node group architecture — significant investment in ASG design, compliance tooling
  3. Regulatory constraints — some frameworks require ASG-backed provisioning audit trails
  4. Cluster API / bare metal — CA is the only mature option
  5. Team familiarity and working-well CA deployment — migration cost may not justify benefit

When Karpenter Is the Right Choice

  1. AWS-native, cost optimization priority — right-sizing + consolidation = meaningful cost reduction
  2. Diverse and variable workloads — batch, spot, GPU, stateless APIs — Karpenter handles all with a few NodePools
  3. Spot-heavy clusters — native interruption handling, diversification, no NTH
  4. Declarative infrastructure-as-code culture — NodePools version cleanly in Git
  5. Low-latency scaling requirements — event-driven workloads, KEDA-triggered jobs, sharp traffic spikes

Running Both: Migration Path and Gotchas

Separating Responsibility

Use labels and taints to prevent CA and Karpenter from managing the same nodes:
# NodePool with taint — CA-managed pods won't tolerate this
spec:
  template:
    metadata:
      labels:
        provisioner: karpenter
    spec:
      taints:
        - key: karpenter.sh/provisioned
          value: "true"
          effect: NoSchedule

Gradual Migration

  1. Phase 1 — Karpenter manages spot/batch workloads. CA manages on-demand production nodes.
  2. Phase 2 — Migrate spot workloads fully. Remove AWS NTH.
  3. Phase 3 — Migrate on-demand. Reduce CA node group capacity gradually.
  4. Phase 4 — Decommission CA once all groups are empty.

Key Gotchas

  • Karpenter consolidation + permissive PDBs — maxUnavailable: 100% will cause disruptive consolidation. Audit PDBs before enabling WhenEmptyOrUnderutilized.
  • NodePool limits are hard stops — pods go Pending indefinitely at limit. Monitor utilization.
  • AMI drift — @latest alias picks up new AMIs on new nodes. Consider pinning for strict change control.
  • Simultaneous scale-down conflicts — use strict label/taint segregation during migration.

Karpenter vs Cluster Autoscaler Cost: When the Savings Are Real

The number most often quoted for “Karpenter vs Cluster Autoscaler cost” is a 20-40% reduction in compute spend after migration. It is a real number — but it comes from a specific starting point, and it is worth being precise about where the money comes from, because two of the three sources are available to Cluster Autoscaler too. Where Karpenter’s savings actually come from:
  • Instance selection across the whole catalog. A pending pod that needs 3 vCPU and 12 GiB lands on the cheapest instance that fits it — possibly a type you would never have created a node group for. With CA you pay for the granularity of the node groups you happened to define. This is the source that CA structurally cannot match, and in heterogeneous clusters it is the largest one.
  • Active consolidation. Karpenter continuously replaces an underused large node with a smaller one that still fits every pod. CA’s scale-down is binary — a node is either below the utilisation threshold and removed, or left alone — so a cluster of 60%-utilised nodes never gets cheaper under CA and keeps shrinking under Karpenter.
  • Spot without ceremony. Diversified spot requests, interruption handling and replacement provisioning are built in. CA can run spot node groups, but the diversification and the graceful replacement are on you (or on the Node Termination Handler).
Where Cluster Autoscaler is just as cheap:
  • Homogeneous, reserved capacity. If your fleet is 90% one instance family covered by Savings Plans or Reserved Instances, the cheapest node is the one you already committed to. Karpenter’s catalog-wide bin packing has nothing to optimise, and its consolidation can actually hurt by moving load off committed capacity onto on-demand instances unless you constrain the NodePool tightly.
  • Steady-state workloads. Consolidation saves money on clusters whose shape changes. A cluster that runs the same 40 pods all day at stable requests has no fragmentation to recover.
  • Well-tuned least-waste with a small set of right-sized node groups. Most of the overprovisioning CA gets blamed for is a node-group design problem — three sizes per family, least-waste as the expander and --scale-down-utilization-threshold raised from the default 0.5 close most of the gap on simple clusters.
The honest rule of thumb: the more heterogeneous and bursty the workloads, the larger the gap in Karpenter’s favour. Batch, CI runners, spot-tolerant services and mixed GPU/CPU fleets are where the 30% figures come from. A stable microservice platform on reserved instances should expect single digits, and the migration effort described below may not pay for itself on cost alone — do it for the provisioning latency and the operational model, not the invoice.

Decision Framework

FactorCluster AutoscalerKarpenter
Cloud supportAll clouds + on-premAWS (GA), Azure (stable), GCP (beta)
Provisioning speed4–8 minutes60–120 seconds
Instance flexibilityNode group pre-config requiredFull catalog, runtime selection
Cost optimizationScale-down onlyScale-down + consolidation + right-sizing
Spot integrationVia ASG + NTHNative, first-class
Operational complexityLowerModerate
Cluster API / bare metalYesNo
ConsolidationNoYes
Running on AWS?
├── No → Azure? → Karpenter (stable) or CA
│        GCP?   → CA or GKE NAP (preferred)
│        Other  → Cluster Autoscaler
│
└── Yes → Hard regulatory constraints on non-ASG provisioning?
          ├── Yes → Cluster Autoscaler
          └── No → Cost optimization priority or diverse workloads?
                   ├── Yes → Karpenter
                   └── No → Either (flip for team preference)

Frequently Asked Questions

Is Karpenter a drop-in replacement for Cluster Autoscaler?

No. Different configuration model, different concepts. Migration requires re-expressing node group config as NodePools/NodeClasses, auditing PDBs, and running both in parallel. Budget at least a sprint for a medium-sized cluster.

Can I run Karpenter on self-managed Kubernetes (not EKS)?

Yes, but non-trivial. Karpenter requires IAM credentials (IRSA or equivalent) to call EC2 APIs. On self-managed clusters, this requires more setup than on EKS where IRSA is built-in.

How does Karpenter interact with HPA and VPA?

No conflict. HPA creates pods u2192 pods go Pending if insufficient nodes u2192 Karpenter provisions nodes u2192 pods scheduled. VPA adjusts pod resource requests, which Karpenter uses as inputs for bin packing.

What happens when Karpenter itself goes down?

Existing nodes and pods continue normally. New pods requiring provisioning go Pending until Karpenter recovers. Scale-down and consolidation pause. Deploy multiple replicas with leader election for production.

Does Karpenter support GPU nodes?

Yes. GPU instance types (p3, p4, g4, g5) can be included in NodePool requirements. Create dedicated NodePools with appropriate taints for GPU-requesting pods.

How does Karpenter handle AMI updates?

The expireAfter field forces node rotation. When a node expires, Karpenter pre-provisions a replacement with the latest AMI per EC2NodeClass, then drains and terminates the old node — a rolling AMI update mechanism without additional tooling.

Is Cluster Autoscaler still actively maintained?

Yes. CA remains under active development in kubernetes/autoscaler, with releases tracking Kubernetes minor versions. It is not being deprecated. For non-AWS environments and working CA deployments, it remains a fully supported and rational choice.

Is Karpenter cheaper than Cluster Autoscaler?

Usually, but not always. Karpenter saves money through three mechanisms: choosing the cheapest instance from the whole catalog for each pending workload, continuously consolidating underused nodes into smaller ones, and handling spot diversification and interruptions natively. On heterogeneous, bursty clusters this adds up to 20-40% less compute spend. On a homogeneous fleet covered by Reserved Instances or Savings Plans, with steady workloads and well-designed node groups using the least-waste expander, Cluster Autoscaler is just as cheap and the migration rarely pays for itself on cost alone.

Karpenter vs Cluster Autoscaler: Quick Decision Recap

Searching for Karpenter vs Cluster Autoscaler? The one-paragraph version: choose Karpenter on AWS when provisioning speed (60–90 seconds vs 4–8 minutes), instance-type flexibility and cost consolidation matter. Choose Cluster Autoscaler when you need multi-cloud portability, you already operate stable node groups, or your platform depends on the node-group model (compliance-pinned AMIs, strict capacity planning).

Related reading: GPU scheduling with Dynamic Resource Allocation.

Related reading: Kamal vs Kubernetes.

Related reading: EKS Auto Mode.

Both can coexist during a migration: run Cluster Autoscaler for existing node groups while Karpenter handles new dynamic capacity, then consolidate once you trust the disruption behavior.


Tested against Kubernetes 1.31–1.36. Karpenter v1.x API (GA; v1.14.1 LTS is current). Cluster Autoscaler 1.36.x. AWS provider examples; Azure and GCP provider details may differ. Last reviewed September 2026.

Kubernetes Resource Requests and Limits: The Complete Production Guide

Kubernetes Resource Requests and Limits: The Complete Production Guide
Your pods are being OOMKilled at 3 AM. Your latency p99 spikes every few minutes with no obvious cause. Your cluster scheduler is placing workloads on nodes that can’t sustain them. In most production Kubernetes incidents, misconfigured resource requests and limits are either the direct cause or an accelerating factor. This is not a “what are requests and limits” tutorial. It is a deep technical guide for engineers who run Kubernetes in production and need to understand what actually happens inside the kernel when these values are set — and what the consequences are when they are wrong.

What Requests and Limits Actually Are

The Kubernetes documentation explains requests and limits at the API level. What it underexplains is the enforcement mechanism: cgroups. When the kubelet admits a pod onto a node, it creates a cgroup hierarchy for that pod under /sys/fs/cgroup/. Each container in the pod gets its own cgroup. The values you set in your pod spec translate directly into cgroup parameters: CPU request → cpu.shares (cgroups v1) or cpu.weight (cgroups v2) CPU limit → cpu.cfs_quota_us and cpu.cfs_period_us Memory request → memory.soft_limit_in_bytes (advisory, used for eviction scoring) Memory limit → memory.limit_in_bytes (hard enforcement, triggers OOMKill) The scheduler uses requests to make placement decisions. It does not know about actual utilization — it knows about committed capacity. A node with 4 cores where running pods have a total CPU request of 3.5 cores has 0.5 cores of schedulable capacity remaining, even if actual CPU utilization is 15%. This is why you can have a fully “utilized” cluster (by requests) where nodes are idle, and why you can have nodes at 95% CPU utilization that still accept new pods because their requests are low. The kubelet uses limits to enforce runtime constraints via those cgroup parameters. The scheduler never sees limits.

CPU vs Memory: Why They Behave Fundamentally Differently

This is the most consequential thing to understand about Kubernetes resource management, and it is routinely misunderstood even by experienced engineers.

CPU Is Compressible

CPU is a time-shared resource. If your container tries to use more CPU than its limit allows, the Linux CFS scheduler simply throttles it — it stops getting CPU time until the next scheduling period. The process continues. It just waits. From the application’s perspective: things slow down. Latency increases. Throughput drops. But the process does not die.

Memory Is Not Compressible

Memory is not time-shared. If your container tries to allocate memory beyond its limit, there is no “slow down” path. The Linux OOM killer selects a process in the cgroup and kills it. The container dies. From the application’s perspective: the process is terminated. Kubernetes restarts the container. You see OOMKilled in kubectl describe pod.
PropertyCPUMemory
EnforcementCFS throttlingOOM Kill
Process survives?Yes (degraded performance)No (killed and restarted)
Compressible?YesNo
Scheduler visibilityRequests onlyRequests only
Over-limit consequenceLatency spikesContainer restart
Setting limits: recommended?Situational (see below)Always
This asymmetry drives every recommendation in the rest of this guide.

QoS Classes: Eviction Priority Under Pressure

Kubernetes assigns each pod a Quality of Service (QoS) class based on the requests and limits set across all its containers. This class determines eviction priority when a node is under memory pressure.

Guaranteed

Condition: Every container has CPU and memory requests and limits set, and requests equal limits for both CPU and memory.
resources:
  requests:
    cpu: "500m"
    memory: "512Mi"
  limits:
    cpu: "500m"
    memory: "512Mi"
Guaranteed pods are the last to be evicted. The kubelet will exhaust BestEffort and Burstable pods before touching these. They get the most predictable resource allocation on the node. Warning: Guaranteed does not mean “always available.” It means “last to be killed.” On a heavily overloaded node, even Guaranteed pods can be evicted.

Burstable

Condition: At least one container has a CPU or memory request or limit set, but the pod does not meet Guaranteed criteria.
resources:
  requests:
    cpu: "250m"
    memory: "256Mi"
  limits:
    cpu: "1000m"
    memory: "1Gi"
Burstable pods are evicted after BestEffort but before Guaranteed. They can burst above their request when capacity is available, but they are not protected when the node is under pressure.

BestEffort

Condition: No container in the pod has any CPU or memory requests or limits set.
# No resources block at all
BestEffort pods are evicted first, always. They get whatever capacity is left over after scheduled workloads consume their requested share. On a loaded node, they may be starved entirely. In production: never run stateful workloads or business-critical services as BestEffort. The Kubernetes scheduler will place them anywhere, and the kubelet will kill them first.

Common Misconfiguration Patterns and Their Consequences

Pattern 1: No Requests or Limits Set

Effect: BestEffort QoS. First to be evicted under memory pressure. Scheduler places pods arbitrarily — it has no data for placement decisions, so it defaults to LeastRequestedPriority, which effectively means these pods may land on the same nodes as heavily-loaded workloads. Real consequence: Your “lightweight” background jobs kill your API servers at 3 AM when a memory spike triggers eviction and BestEffort pods happen to be sitting next to them on the same node.

Pattern 2: Requests Equal Limits (Guaranteed QoS)

This is the common “safe” pattern recommended in older Kubernetes documentation. It is not wrong, but it has a trap: CPU limits = CPU requests means CPU throttling is guaranteed to trigger. Your pod will be throttled the moment it tries to burst above the request — during startup, during GC, during a traffic spike — even if the node has abundant free CPU. For latency-sensitive applications, this means predictable throttling spikes at exactly the moments you need the most CPU. Memory: Setting memory request = memory limit is appropriate and recommended. The behavior is correct: the pod runs in a controlled memory budget.

Pattern 3: Limits Much Higher Than Requests (Burstable with High Ratio)

resources:
  requests:
    cpu: "100m"
    memory: "128Mi"
  limits:
    cpu: "4000m"
    memory: "4Gi"
This is the opposite extreme. The scheduler thinks this pod needs 100m CPU and 128Mi memory. Dozens of these can be scheduled onto a single node. When they all burst simultaneously — which they will, during a deployment, a traffic event, or a GC cycle — the node is overloaded, memory pressure triggers OOMKill cascades, and the scheduler has no idea anything is wrong because the committed capacity (by requests) looks fine. The limit:request ratio matters. A 10x or 20x memory limit:request ratio on many pods is a recipe for node instability. A reasonable starting point is 2x–4x for memory, less for CPU.

Pattern 4: CPU Limits Set to “Be Safe”

This is the subtlest misconfiguration and the one with the most hidden latency impact. We cover it in depth in the next section.

The CPU Throttling Problem: CFS Bandwidth and Hidden Latency

This is where many production Kubernetes deployments have a silent performance problem they cannot easily diagnose.

How CFS Bandwidth Throttling Works

The Linux Completely Fair Scheduler (CFS) enforces CPU limits using bandwidth control. The relevant parameters are:
  • cpu.cfs_period_us: the accounting period, default 100ms
  • cpu.cfs_quota_us: how many microseconds of CPU time the cgroup can use per period
If you set cpu: "500m" as a limit, Kubernetes sets cpu.cfs_quota_us = 50000 (50ms per 100ms period). This means the container can use at most 50% of one CPU core per 100ms window. The problem: quota is enforced per period, not as a moving average. If your container uses its full 50ms allocation in the first 60ms of a period, it is throttled for the remaining 40ms — even if the node has 7 idle CPUs. The CPU sits idle. Your container waits.

Why This Causes Latency Spikes Even at Low Utilization

This is counterintuitive and the source of many production mysteries. You can have a container running at 10% average CPU utilization that is regularly throttled, because its instantaneous CPU usage within a single 100ms window exceeds its quota. Java applications with JVM garbage collection are particularly vulnerable. GC causes a CPU burst of short duration. If that burst exceeds the per-period quota, the GC pause is extended artificially by throttling — even though the GC event itself would have been short. The same applies to Node.js event loop processing, Python import at startup, and any application that has bursty CPU behavior (which is most of them).

The Cloudflare and Netflix Evidence

Cloudflare published findings showing that CPU throttling was responsible for significant tail latency increases in their containerized workloads, and that removing CPU limits reduced p99 latency substantially for services that appeared to have headroom. Netflix has documented similar patterns in their capacity planning work, noting that per-period quota enforcement does not model real application CPU behavior accurately. The kernel community has been aware of this for years. The fix — moving to cgroups v2 with better scheduler integration — helps but does not eliminate the problem. Kubernetes 1.25+ with cgroups v2 nodes experience less throttling under the same limits, but the fundamental issue remains: CPU limits throttle bursty applications unpredictably.

The Recommendation: Consider Not Setting CPU Limits

This is controversial but grounded in the evidence: For latency-sensitive services: do not set CPU limits. Set CPU requests accurately and rely on the scheduler for placement. The argument: – CPU throttling is a soft failure mode that is hard to observe and diagnose – OOMKill is a hard failure mode that is visible and recoverable – CPU requests give the scheduler accurate placement data without creating throttling – Nodes handle CPU oversubscription gracefully through time-sharing; they do not handle memory oversubscription gracefully When to still set CPU limits: – Multi-tenant clusters where noisy neighbor isolation is critical – Batch workloads where predictable CPU allocation matters more than latency – When your monitoring and alerting can catch CPU starvation at the node level When you do not set CPU limits, you must set CPU requests accurately. A request of 100m for a service that normally uses 800m means the scheduler places it on a node that cannot actually sustain it. The result is real CPU starvation, not artificial throttling — but it is CPU starvation nonetheless.

Memory: Always Set Limits

The contrast with CPU is direct. Memory is non-compressible. A container that leaks memory or has a runaway allocation will consume all available node memory if unconstrained. This does not degrade gracefully — it triggers the OOM killer, which may kill unrelated processes on the node. Always set memory limits. Always. The consequence — OOMKill — is visible, logged, and Kubernetes handles it by restarting the container. An OOMKilled exit code is actionable: you either have a memory leak, your limit is too low, or your sizing methodology is wrong. All three are diagnosable. The alternative — no memory limit — means a single leaking pod can destabilize an entire node and trigger eviction cascades affecting unrelated workloads. Set memory requests equal to the p95 steady-state usage of your application. Set memory limits at 1.5x–2x the request to absorb traffic spikes and GC pressure. Profile your application under load to establish these baselines.

Vertical Pod Autoscaler (VPA)

VPA is the Kubernetes component designed to solve the sizing problem automatically. It observes actual resource utilization and recommends (or applies) adjusted requests.

How VPA Works

VPA has three components:
  • Recommender: Watches historical metrics and computes recommended requests based on observed utilization. Does not modify pods.
  • Updater: Evicts pods whose current requests differ significantly from recommendations (when VPA mode is Auto or Recreate).
  • Admission Controller: Mutates pod specs at admission time to apply recommendations from the Recommender.
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
  name: api-server-vpa
  namespace: production
spec:
  targetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: api-server
  updatePolicy:
    updateMode: "Off"   # Recommend only — do not evict pods
  resourcePolicy:
    containerPolicies:
    - containerName: api-server
      minAllowed:
        cpu: 100m
        memory: 128Mi
      maxAllowed:
        cpu: 4000m
        memory: 4Gi
      controlledResources: ["cpu", "memory"]
      controlledValues: RequestsAndLimits

VPA Modes

ModeBehavior
OffCompute recommendations only. No pod mutations.
InitialApply recommendations to new pods only. Do not evict running pods.
RecreateEvict pods when recommendations change significantly.
AutoCurrently equivalent to Recreate. May change in future versions.

When to Use VPA

Right-sizing during initial rollout: Run VPA in Off mode for 1–2 weeks on a new service. Review recommendations before applying. This is the most valuable use case. Services with unpredictable or seasonal load patterns: VPA adapts requests based on observed behavior. Combined with HPA for horizontal scaling, this gives you right-sized replicas that scale out horizontally. VPA and HPA cannot both manage the same metric. If HPA is scaling on CPU utilization, do not use VPA with controlledValues: RequestsAndLimits for CPU — they will fight each other. Use controlledValues: RequestsOnly and let HPA manage scale. VPA limitations: – Requires pod restarts to apply recommendations (Updater evicts pods) – Does not work well with stateful workloads in strict availability windows – Recommender needs sufficient history (at least a few days) to produce reliable recommendations – Does not account for traffic spikes that haven’t been observed yet

LimitRange and ResourceQuota: Namespace-Level Guardrails

Requests and limits on individual pods solve the per-workload problem. LimitRange and ResourceQuota solve the namespace and cluster-level governance problem.

LimitRange

LimitRange sets default requests and limits for containers in a namespace, and enforces minimum/maximum boundaries. Any pod admitted to the namespace that does not have explicit requests/limits set will receive the defaults.
apiVersion: v1
kind: LimitRange
metadata:
  name: default-limits
  namespace: production
spec:
  limits:
  - type: Container
    default:
      cpu: "500m"
      memory: "512Mi"
    defaultRequest:
      cpu: "100m"
      memory: "128Mi"
    max:
      cpu: "4000m"
      memory: "8Gi"
    min:
      cpu: "50m"
      memory: "64Mi"
  - type: Pod
    max:
      cpu: "8000m"
      memory: "16Gi"
  - type: PersistentVolumeClaim
    max:
      storage: "50Gi"
    min:
      storage: "1Gi"
Key behaviors: – default applies as the limit for containers that set a request but no limit – defaultRequest applies as the request for containers that set no request – max and min cause admission to fail if violated – LimitRange applies at admission time — changing it does not affect running pods Use LimitRange to: – Prevent BestEffort pods from being admitted (by setting defaultRequest values) – Enforce organizational standards for minimum resource specifications – Protect the cluster from pods requesting unbounded resources

ResourceQuota

ResourceQuota limits the total amount of resources that can be consumed by all pods in a namespace. This is the multi-tenant governance tool.
apiVersion: v1
kind: ResourceQuota
metadata:
  name: production-quota
  namespace: production
spec:
  hard:
    requests.cpu: "20"
    requests.memory: "40Gi"
    limits.cpu: "40"
    limits.memory: "80Gi"
    pods: "100"
    persistentvolumeclaims: "20"
    requests.storage: "500Gi"
    count/deployments.apps: "50"
    count/services: "50"
    count/secrets: "100"
    count/configmaps: "100"
Critical interaction with LimitRange: When ResourceQuota is active in a namespace, every pod must have requests and limits set or it will be rejected. This is why LimitRange defaults are important — they ensure pods without explicit resources are not rejected by the quota system. Use ResourceQuota to: – Enforce team/application resource budgets in shared clusters – Prevent runaway deployments from consuming all cluster capacity – Implement chargeback policies (track resource consumption per namespace)

Practical Sizing Methodology

Step 1: Instrument Before You Set Values

Deploy initially with only requests set (no CPU limits, memory limits set conservatively high) and monitor for 2–4 weeks under realistic load. Useful PromQL queries for sizing:
# p95 CPU usage over the last 7 days
histogram_quantile(0.95,
  rate(container_cpu_usage_seconds_total{
    container="api-server",
    namespace="production"
  }[5m])
)

# p99 memory working set over the last 7 days
quantile_over_time(0.99,
  container_memory_working_set_bytes{
    container="api-server",
    namespace="production"
  }[7d]
)

# CPU throttling ratio (alert if >5%)
rate(container_cpu_cfs_throttled_seconds_total{container="api-server"}[5m])
/
rate(container_cpu_cfs_periods_total{container="api-server"}[5m])

Step 2: Set CPU Requests from p95 Observations

Set CPU request = p95 CPU usage under realistic production load. For latency-sensitive services: do not set CPU limits. For batch or background jobs: set CPU limits at 2x–4x the request.

Step 3: Set Memory Requests and Limits

Set memory request = p95 memory working set over at least 7 days. Set memory limit = max(observed peak, 1.5 × request). For Java/Python with large processing, use 2x.
# Production example: Java microservice
resources:
  requests:
    cpu: "500m"       # p95 observed: ~420m
    memory: "768Mi"   # p95 observed: ~680Mi
  limits:
    # No CPU limit — latency-sensitive service
    memory: "1.5Gi"   # 2x request, covers GC pressure

Step 4: Use VPA Recommendations to Validate

Run VPA in Off mode alongside your manually-set values. After 1–2 weeks, compare VPA recommendations to your current settings.

Step 5: Adjust for Workload Lifecycle Events

Account for: JVM warmup at startup (CPU spike 3–10x steady-state), rolling deployment overlap (namespace quota headroom), and known traffic peaks (size to peak, not average).

Decision Framework: What to Set Based on Workload Type

Workload TypeCPU RequestCPU LimitMemory RequestMemory LimitQoS Target
Latency-sensitive API (Go, Java, Node)p95 observedDo not setp95 observed1.5–2x requestBurstable
Batch / background jobsp50 observed2–4x requestp95 observed1.5x requestBurstable
System-critical (coredns, metrics-server)ConservativeEqual to requestConservativeEqual to requestGuaranteed
Stateful / databases (in-cluster)p95 observedDo not setp99 observed1.25x requestBurstable
Dev/test workloadsLow (100m)2x requestLow (128Mi)2x requestBurstable
Sidecar containers (envoy, otel-collector)Profile individuallyContextualProfile individually1.5x requestMatches primary

Monitoring and Alerting

# OOMKill rate
- alert: ContainerOOMKilled
  expr: increase(kube_pod_container_status_last_terminated_reason{reason="OOMKilled"}[5m]) > 0
  for: 0m
  labels:
    severity: warning

# CPU throttling >10%
- alert: CPUThrottlingHigh
  expr: |
    rate(container_cpu_cfs_throttled_seconds_total[5m])
    /
    rate(container_cpu_cfs_periods_total[5m])
    > 0.10
  for: 5m
  labels:
    severity: warning

# Memory near limit >85%
- alert: MemoryNearLimit
  expr: |
    container_memory_working_set_bytes
    /
    (container_spec_memory_limit_bytes > 0)
    > 0.85
  for: 5m
  labels:
    severity: warning

FAQ

Q: My Java application keeps getting OOMKilled but I’ve set limits at 2x average usage. What am I missing?

The JVM heap (-Xmx) is not the only memory consumer. Off-heap buffers, Metaspace, thread stacks, and JVM overhead add 25–40% on top. Set -Xmx at ~75% of your container memory limit. For a 1Gi limit: -Xmx768m is a safe starting point.

Q: Should I set the same resources in all environments?

No. Dev/test can use lower values. But the ratio between request and limit should be similar, and the resource profile should be close enough to catch misconfigurations before production.

Q: Can I use HPA and VPA together?

Yes, carefully. Use HPA for replica scaling (CPU or custom metrics) and VPA in Off mode or controlledValues: RequestsOnly for right-sizing guidance. Never have both managing the same metric simultaneously.

Q: My cluster uses cgroups v2. Does CPU throttling still apply?

Improved but not eliminated. cgroups v2 uses a weight-based scheduler that reduces throttling artifacts. However, cpu.cfs_quota_us enforcement still exists when CPU limits are set. For latency-sensitive workloads, the case for not setting CPU limits remains valid on cgroups v2.

Q: What is a realistic cluster overcommit ratio?

CPU: 5–10x overcommit (total requests vs physical cores) is common for mixed workloads with accurate requests. Memory: 1.5–2x cluster-level overcommit is manageable at 1.5x request:limit ratios. Beyond 2x, node memory pressure events become frequent.

Q: LimitRange is set but pods are still admitted without resources. Why?

LimitRange defaults only apply to containers with no resource specification at all. If a container specifies requests.cpu but not limits.cpu, the LimitRange default for CPU does not fill in the missing limit. Also verify the LimitRange is in the correct namespace: kubectl get limitrange -n <namespace>.

Q: What does a pod with no memory limit do to a node?

It can consume all available node memory unconstrained. This triggers the Linux OOM killer at the node level, which may kill processes outside the container — including the kubelet itself in extreme cases. Memory limits are non-negotiable in production.


Tested against Kubernetes 1.28–1.32. cgroups v2 behavior noted where it differs from v1. VPA examples use autoscaling.k8s.io/v1 API (VPA 0.14+).

Prometheus 3.0 and OpenTelemetry: Native OTLP Support Explained

Prometheus 3.0 and OpenTelemetry: Native OTLP Support Explained

Seven years is a long time in observability. Since Prometheus 2.0 landed in 2017, the ecosystem has been transformed by cloud-native adoption, the rise of distributed tracing, and the emergence of OpenTelemetry as the de facto standard for instrumentation. Prometheus 3.0, released in November 2024, is the project’s answer to that transformation — and its most significant change is the native ability to ingest OpenTelemetry metrics directly, without an intermediary collector standing in the way.

Related reading: Prometheus cardinality and scaling.

This article goes deep on what Prometheus 3.0 actually changes for platform engineers and cloud architects who are running — or planning to run — OTel-instrumented workloads alongside Prometheus-based monitoring stacks. We will cover the native OTLP ingestion endpoint, UTF-8 metric name support, Remote Write 2.0, migration considerations, and the architectural patterns that still make sense even when native OTLP is available.

What Changed in Prometheus 3.0: The OTel-Relevant Picture

Prometheus 3.0 ships a substantial set of changes. Not all of them are equally relevant to OpenTelemetry integration, so let’s focus on what actually moves the needle for OTel users before diving into each area in detail.

Native OTLP Ingestion

The flagship feature: Prometheus 3.0 ships with a built-in OTLP receiver that exposes an HTTP endpoint accepting metrics in the OpenTelemetry Protocol format. Applications instrumented with any OTel SDK can now push metrics directly to Prometheus without routing through an OpenTelemetry Collector. This is not a sidecar, not a plugin, not an external adapter — it is a first-class endpoint in the Prometheus binary itself.

UTF-8 Metric Names

Prometheus historically restricted metric names to [a-zA-Z_:][a-zA-Z0-9_:]*. OpenTelemetry uses dots and slashes in metric names by convention — http.server.request.duration is a canonical OTel metric name. Prometheus 3.0 lifts this restriction and supports arbitrary UTF-8 characters in metric names and label names, which is the single most important compatibility change for OTel interoperability.

Remote Write 2.0

Remote Write 2.0

Remote Write 2.0 replaces the original protocol with a more efficient encoding based on protobuf, adds native histogram support in the wire format, and reduces bandwidth consumption significantly for large-scale deployments. If you are federating metrics to Thanos, Mimir, or Cortex, this matters for operational cost.

New UI

The Prometheus web UI has been completely rewritten. The new UI uses React, supports metric metadata exploration, and provides a significantly improved query-building experience. This is a quality-of-life improvement rather than an architectural change, but it reduces the dependency on external tools like Grafana for ad-hoc investigation.

Breaking Changes Summary

Prometheus 3.0 removes several features that were deprecated in 2.x. The most operationally significant are: removal of the --web.enable-admin-api deprecated flag path, removal of certain legacy storage format options, changes to default scrape timeouts, and stricter validation of configuration that was previously silently accepted. We cover a migration checklist later in this article.

The OTLP Receiver: How It Works and What It Accepts

The OTLP receiver in Prometheus 3.0 is implemented as an optional feature that must be explicitly enabled. Once enabled, it exposes an HTTP endpoint at /api/v1/otlp/v1/metrics that accepts protobuf-encoded OTLP ExportMetricsServiceRequest payloads — the same wire format used by the OpenTelemetry Collector’s OTLP exporter.

What It Accepts (and What It Does Not)

This is critical to understand before you architect around native OTLP ingestion: Prometheus 3.0 OTLP support is metrics-only. It does not accept traces or logs. OTLP is a unified protocol covering all three signals, but Prometheus is a metrics store — the receiver handles only the metrics portion of the OTLP specification.

Supported metric types in the OTLP receiver:

  • Gauge — maps directly to a Prometheus Gauge
  • Sum (monotonic) — maps to a Prometheus Counter
  • Sum (non-monotonic) — maps to a Prometheus Gauge
  • Histogram (explicit bucket) — maps to a Prometheus Histogram
  • ExponentialHistogram — maps to Prometheus Native Histograms (a 3.0 feature)
  • Summary — maps to a Prometheus Summary

Resource attributes from the OTLP payload — things like service.name, k8s.pod.name, cloud.region — are converted to Prometheus labels. This conversion is configurable, and by default Prometheus applies a promotion strategy that converts the most common resource attributes to labels while discarding ones that would create extremely high cardinality.

Enabling the OTLP Receiver

Enabling native OTLP ingestion requires two things: a feature flag and a configuration block in prometheus.yml.

Start the Prometheus binary with the feature flag:

prometheus 
  --config.file=/etc/prometheus/prometheus.yml 
  --enable-feature=otlp-write-receiver

Then add the OTLP receiver configuration to your prometheus.yml:

# prometheus.yml
global:
  scrape_interval: 15s
  evaluation_interval: 15s

otlp:
  # Promote these OTLP resource attributes to Prometheus labels
  promote_resource_attributes:
    - service.name
    - service.namespace
    - service.instance.id
    - k8s.namespace.name
    - k8s.pod.name
    - k8s.node.name
    - cloud.region
    - deployment.environment

With this configuration, Prometheus will listen on port 9090 (default) and accept OTLP metrics at http://<prometheus-host>:9090/api/v1/otlp/v1/metrics.

Resource Attribute Promotion Strategy

The promote_resource_attributes list deserves careful thought. OTLP carries rich resource-level context — every metric payload includes a ResourceMetrics object with attributes describing the source: service name, version, environment, Kubernetes pod, node, cluster, cloud provider details, and more. Prometheus labels are flat key-value pairs on each time series. Promoting too many resource attributes explodes cardinality; promoting too few loses important context.

A pragmatic starting list for Kubernetes deployments:

otlp:
  promote_resource_attributes:
    - service.name          # Critical: identifies the service
    - service.namespace     # Logical grouping
    - deployment.environment  # prod/staging/dev
    - k8s.namespace.name    # Kubernetes namespace
    - k8s.pod.name          # Pod-level cardinality — consider omitting in high-scale
    - k8s.node.name         # Useful for infrastructure correlation

Avoid blindly promoting k8s.pod.name at scale — in a cluster with thousands of short-lived pods, this creates significant cardinality pressure. Prefer service.name and service.namespace for most alerting use cases, reserving pod-level labels for debugging dashboards.

UTF-8 Metric Names: Why This Is the Real Game-Changer

To appreciate why UTF-8 metric name support matters so much, you need to understand the friction it eliminates. OpenTelemetry semantic conventions define metric names using dots as namespace separators. The canonical HTTP server duration metric is http.server.request.duration. The canonical database query duration is db.client.operation.duration. These names are standardized across languages and frameworks — your Go service and your Java service and your Python service all emit the same metric name when instrumented with OTel.

Prometheus 2.x could not store these names. The dots are illegal characters in Prometheus metric naming. Every OTel-to-Prometheus bridge — the OpenTelemetry Collector’s Prometheus exporter, prom-client compatibility layers, the older prometheusremotewrite exporter — had to translate these names, typically by replacing dots with underscores: http_server_request_duration.

This translation is lossy and creates multiple problems:

  • Name collisions: http.server.request_duration and http.server.request.duration both become http_server_request_duration
  • Dashboard breakage: Grafana dashboards built against OTel semantic conventions don’t work against translated Prometheus metrics without modification
  • Cross-signal correlation: Trace attributes use dot notation; when metric names differ, automated correlation tools lose the thread
  • Vendor lock-in pressure: Teams end up with separate naming conventions for “Prometheus metrics” vs “OTel metrics” and maintain both

Prometheus 3.0 with UTF-8 support stores http.server.request.duration natively. No translation. No collision. The metric name you instrument with is the metric name you query.

Enabling UTF-8 Metric Names

UTF-8 metric names require the utf8-names feature flag:

prometheus 
  --config.file=/etc/prometheus/prometheus.yml 
  --enable-feature=utf8-names 
  --enable-feature=otlp-write-receiver

Once enabled, PromQL queries must use quoted metric names when the name contains characters outside the legacy character set:

# Legacy metric name — unquoted works fine
http_server_requests_total

# OTel metric name with dots — requires quoting in PromQL
{"__name__"="http.server.request.duration"}

# Or using the new PromQL syntax in Prometheus 3.0
http.server.request.duration{service_name="api-gateway"}

The PromQL parser in Prometheus 3.0 has been updated to handle quoted metric names as a first-class construct. Grafana’s PromQL engine has also been updated to handle this syntax — verify your Grafana version (10.3+ has full support) before deploying.

OTel SDK to Prometheus 3.0 Directly: No Collector Required

For teams that only need to get application metrics into Prometheus, native OTLP ingestion enables a dramatically simpler architecture. Here’s what it looks like with different OTel SDKs.

Go (OpenTelemetry SDK)

package main

import (
    "context"
    "time"

    "go.opentelemetry.io/otel"
    "go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetrichttp"
    "go.opentelemetry.io/otel/sdk/metric"
    "go.opentelemetry.io/otel/sdk/resource"
    semconv "go.opentelemetry.io/otel/semconv/v1.26.0"
)

func initMetrics(ctx context.Context) (*metric.MeterProvider, error) {
    res, err := resource.New(ctx,
        resource.WithAttributes(
            semconv.ServiceName("my-api"),
            semconv.ServiceNamespace("platform"),
            semconv.DeploymentEnvironment("production"),
        ),
    )
    if err != nil {
        return nil, err
    }

    // Point directly at Prometheus 3.0 OTLP endpoint
    exporter, err := otlpmetrichttp.New(ctx,
        otlpmetrichttp.WithEndpoint("prometheus:9090"),
        otlpmetrichttp.WithURLPath("/api/v1/otlp/v1/metrics"),
        otlpmetrichttp.WithInsecure(), // Use WithTLSClientConfig for production
    )
    if err != nil {
        return nil, err
    }

    provider := metric.NewMeterProvider(
        metric.WithResource(res),
        metric.WithReader(
            metric.NewPeriodicReader(exporter,
                metric.WithInterval(30*time.Second),
            ),
        ),
    )

    otel.SetMeterProvider(provider)
    return provider, nil
}

Python (OpenTelemetry SDK)

from opentelemetry import metrics
from opentelemetry.sdk.metrics import MeterProvider
from opentelemetry.sdk.metrics.export import PeriodicExportingMetricReader
from opentelemetry.exporter.otlp.proto.http.metric_exporter import OTLPMetricExporter
from opentelemetry.sdk.resources import Resource, SERVICE_NAME, SERVICE_NAMESPACE

resource = Resource.create({
    SERVICE_NAME: "my-api",
    SERVICE_NAMESPACE: "platform",
    "deployment.environment": "production",
})

exporter = OTLPMetricExporter(
    endpoint="http://prometheus:9090/api/v1/otlp/v1/metrics",
)

reader = PeriodicExportingMetricReader(
    exporter,
    export_interval_millis=30_000,
)

provider = MeterProvider(resource=resource, metric_readers=[reader])
metrics.set_meter_provider(provider)

# Use the meter
meter = metrics.get_meter("my-api")
request_counter = meter.create_counter(
    name="http.server.request.count",
    description="Total HTTP server requests",
    unit="1",
)
request_duration = meter.create_histogram(
    name="http.server.request.duration",
    description="HTTP server request duration",
    unit="s",
)

Java (OpenTelemetry SDK with Spring Boot)

# application.properties (Spring Boot with OTel auto-instrumentation)
otel.service.name=my-api
otel.resource.attributes=service.namespace=platform,deployment.environment=production

# Configure OTLP exporter to push directly to Prometheus
otel.metrics.exporter=otlp
otel.exporter.otlp.metrics.endpoint=http://prometheus:9090/api/v1/otlp/v1/metrics
otel.exporter.otlp.metrics.protocol=http/protobuf

# Export interval
otel.metric.export.interval=30000

With Spring Boot and the OTel Java agent, no code changes are required beyond configuration — the agent instruments your HTTP server, database clients, and messaging systems automatically and pushes metrics using the names defined in OTel semantic conventions.

OTel Collector to Prometheus 3.0: When You Need the Intermediary

Native OTLP ingestion is compelling, but the OpenTelemetry Collector remains relevant for a significant set of use cases. Understanding when each pattern is appropriate is the core architectural decision you will face when adopting Prometheus 3.0 in an OTel environment.

Pattern 1: OTel Collector as Fan-Out Gateway

When you need to send metrics to multiple backends simultaneously — Prometheus for alerting, a long-term store like Thanos for historical analysis, and a commercial observability platform for full-stack correlation — the OTel Collector handles fan-out efficiently. Applications push once to the Collector; the Collector distributes to all backends.

# otel-collector-config.yaml
receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 0.0.0.0:4317
      http:
        endpoint: 0.0.0.0:4318

processors:
  batch:
    timeout: 10s
    send_batch_size: 1000
  memory_limiter:
    check_interval: 1s
    limit_mib: 512

exporters:
  # Push to Prometheus 3.0 via OTLP
  otlphttp/prometheus:
    endpoint: http://prometheus:9090/api/v1/otlp
    tls:
      insecure: true

  # Fan-out to Thanos via remote_write
  prometheusremotewrite/thanos:
    endpoint: http://thanos-receive:10908/api/v1/receive
    resource_to_telemetry_conversion:
      enabled: true

  # Fan-out to commercial backend
  otlp/datadog:
    endpoint: https://otel-intake.datadoghq.com
    headers:
      DD-API-KEY: "${DD_API_KEY}"

service:
  pipelines:
    metrics:
      receivers: [otlp]
      processors: [memory_limiter, batch]
      exporters: [otlphttp/prometheus, prometheusremotewrite/thanos, otlp/datadog]

Pattern 2: Collector for Metric Transformation

The OTel Collector’s transform processor and metricstransform processor allow you to reshape metrics before they reach Prometheus: rename labels, add static attributes, filter out high-cardinality series, aggregate metrics to reduce storage cost, or apply unit conversions. These operations are not available in Prometheus’s native OTLP receiver.

processors:
  transform/metrics:
    metric_statements:
      - context: metric
        statements:
          # Drop internal debug metrics
          - delete_matching_keys(attributes, "internal.*")
          # Normalize environment label values
          - set(attributes["deployment.environment"], "prod")
            where attributes["deployment.environment"] == "production"

  filter/drop_debug:
    metrics:
      exclude:
        match_type: regexp
        metric_names:
          - ".*.debug..*"
          - "runtime.go.internal..*"

  metricstransform:
    transforms:
      # Rename a metric to match your existing Prometheus naming convention
      - include: http.server.request.duration
        action: update
        new_name: http_server_request_duration_seconds

Pattern 3: Collector for Traces and Logs (Always Required)

If your architecture includes traces and logs alongside metrics — and in 2025 it almost certainly does — you need an OTel Collector regardless of what you do with metrics. Prometheus does not accept traces or logs. Jaeger, Tempo, and Loki all have their own ingestion protocols. The Collector is the universal routing layer for the three pillars of observability.

In this architecture, it is usually simpler to route all three signals through the Collector and let it push metrics to Prometheus via OTLP or remote_write, rather than splitting metrics to go directly and everything else through the Collector.

When to Use Native OTLP vs. OTel Collector: Decision Framework

ScenarioNative OTLPOTel Collector
Single metrics backend (Prometheus only)PreferredOverkill
Multiple metrics backendsNot sufficientRequired
Traces + Logs in scopeNot applicableRequired
Metric transformation/filtering neededNot supportedRequired
Simple Kubernetes-native deploymentPreferredAdditional complexity
Air-gapped / constrained environmentsPreferred (fewer components)Consider carefully
Mixed OTel + legacy Prometheus targetsWorks alongside scrapingCan normalize naming
High-volume, need batching/bufferingLimited controlPreferred

The pragmatic recommendation for most platform engineering teams: if you are already running the OTel Collector (and you should be if traces are in scope), continue routing metrics through it. Use the Collector’s otlphttp exporter to push to Prometheus 3.0. Reserve the direct SDK-to-Prometheus pattern for simple services where the Collector would be the only reason to add complexity.

Remote Write 2.0: What Changes for Existing Setups

Remote Write 2.0 is a significant protocol upgrade with real operational implications for teams using Prometheus as a metrics source for long-term storage systems like Thanos, Mimir, VictoriaMetrics, or Cortex.

Key Protocol Changes

  • Protobuf encoding with snappy compression — replacing the previous text-based format. Typically 50-70% reduction in wire size for large metric batches
  • Native histogram support in the wire format — exponential histograms can now be forwarded without converting to classic histograms, preserving full resolution
  • Metadata forwarding — metric type and unit information is now transmitted alongside samples, enabling better downstream processing
  • Created timestamps — the timestamp at which a counter was created is forwarded, enabling more accurate rate calculations across restarts

Configuring Remote Write 2.0

# prometheus.yml
remote_write:
  - url: "http://thanos-receive:10908/api/v1/receive"
    # Remote Write 2.0 is negotiated automatically with compatible receivers
    # Force RW2.0 explicitly if needed:
    send_native_histograms: true
    metadata_config:
      send: true
      send_interval: 1m
    queue_config:
      capacity: 10000
      max_shards: 200
      max_samples_per_send: 2000
      batch_send_deadline: 5s

Remote Write 2.0 uses protocol content negotiation — Prometheus 3.0 will attempt RW2.0 first and fall back to RW1.0 if the receiver does not support it. This means upgrades are generally backward-compatible. Verify that your receiving system (Thanos Receive 0.35+, Mimir 2.12+, VictoriaMetrics 1.98+) supports RW2.0 before expecting the benefits.

Migration from Prometheus 2.x: Breaking Changes Checklist

Upgrading from Prometheus 2.x to 3.0 requires attention to several breaking changes. This checklist covers the operationally significant ones for teams running production Prometheus deployments.

Configuration Changes

  • Removed: query.lookback-delta default change — the default changed from 5 minutes to match the scrape interval. Queries that relied on the 5m default may return different results. Audit alerting rules that use instant queries on counters.
  • Removed: deprecated remote_write options — remote_write[].queue_config.capacity semantics changed. Review and update queue configurations.
  • Removed: storage.tsdb.allow-overlapping-blocks flag — overlapping blocks handling is now automatic. Remove this flag from your startup scripts.
  • Scrape protocols default change — Prometheus 3.0 defaults to OpenMetrics format for scraping when targets support it. This enables native histograms but may surface parsing differences. Test with --enable-feature=no-default-scrape-port removed if you relied on the old behavior.
  • Agent mode changes — if using Prometheus Agent mode, review the updated configuration options for WAL management.

PromQL Changes

  • Stricter parsing — some previously accepted but technically invalid PromQL expressions now fail. Run your alerting rules through promtool check rules against a Prometheus 3.0 binary before cutover.
  • Native histogram functions — new functions like histogram_fraction() and histogram_quantile() have updated behavior with native histograms. Existing dashboard queries using histogram_quantile() on classic histograms continue to work unchanged.

Storage Compatibility

Prometheus 3.0 can read existing 2.x TSDB data. The upgrade path does not require a data migration. However, Prometheus 2.x cannot read data blocks written by 3.0 (downgrade is not supported without data loss after any writes have occurred). Take a snapshot before upgrading if you need rollback capability:

# Take a TSDB snapshot before upgrading
curl -X POST http://prometheus:9090/api/v1/admin/tsdb/snapshot

# Verify the snapshot exists
ls /prometheus/snapshots/

Pre-Upgrade Validation Steps

# 1. Validate configuration against Prometheus 3.0
docker run --rm -v $(pwd)/prometheus.yml:/etc/prometheus/prometheus.yml 
  prom/prometheus:v3.0.0 
  promtool check config /etc/prometheus/prometheus.yml

# 2. Validate alerting rules
docker run --rm -v $(pwd)/rules:/etc/prometheus/rules 
  prom/prometheus:v3.0.0 
  promtool check rules /etc/prometheus/rules/*.yml

# 3. Run in parallel (shadow mode) before full cutover
# Deploy Prometheus 3.0 alongside 2.x, scraping the same targets
# Compare query results between versions using promtool query range

Practical Kubernetes Deployment Example

Here is a production-ready Kubernetes deployment of Prometheus 3.0 with OTLP ingestion enabled, suitable as a starting point for platform engineering teams.

Prometheus 3.0 ConfigMap

apiVersion: v1
kind: ConfigMap
metadata:
  name: prometheus-config
  namespace: monitoring
data:
  prometheus.yml: |
    global:
      scrape_interval: 15s
      evaluation_interval: 15s
      external_labels:
        cluster: production
        region: eu-west-1

    otlp:
      promote_resource_attributes:
        - service.name
        - service.namespace
        - deployment.environment
        - k8s.namespace.name
        - k8s.pod.name

    rule_files:
      - /etc/prometheus/rules/*.yml

    scrape_configs:
      - job_name: kubernetes-pods
        kubernetes_sd_configs:
          - role: pod
        relabel_configs:
          - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape]
            action: keep
            regex: "true"
          - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_path]
            action: replace
            target_label: __metrics_path__
            regex: (.+)

    remote_write:
      - url: http://thanos-receive.monitoring.svc.cluster.local:10908/api/v1/receive
        send_native_histograms: true
        metadata_config:
          send: true

Prometheus 3.0 Deployment

apiVersion: apps/v1
kind: Deployment
metadata:
  name: prometheus
  namespace: monitoring
spec:
  replicas: 1
  selector:
    matchLabels:
      app: prometheus
  template:
    metadata:
      labels:
        app: prometheus
    spec:
      serviceAccountName: prometheus
      containers:
        - name: prometheus
          image: prom/prometheus:v3.0.0
          args:
            - --config.file=/etc/prometheus/prometheus.yml
            - --storage.tsdb.path=/prometheus/data
            - --storage.tsdb.retention.time=15d
            - --web.enable-lifecycle
            - --web.enable-admin-api
            - --enable-feature=otlp-write-receiver
            - --enable-feature=utf8-names
            - --enable-feature=native-histograms
          ports:
            - name: http
              containerPort: 9090
              protocol: TCP
          volumeMounts:
            - name: config
              mountPath: /etc/prometheus
            - name: data
              mountPath: /prometheus/data
          resources:
            requests:
              cpu: 500m
              memory: 2Gi
            limits:
              cpu: 2000m
              memory: 8Gi
          livenessProbe:
            httpGet:
              path: /-/healthy
              port: http
            initialDelaySeconds: 30
            periodSeconds: 15
          readinessProbe:
            httpGet:
              path: /-/ready
              port: http
            initialDelaySeconds: 5
            periodSeconds: 5
      volumes:
        - name: config
          configMap:
            name: prometheus-config
        - name: data
          persistentVolumeClaim:
            claimName: prometheus-data
---
apiVersion: v1
kind: Service
metadata:
  name: prometheus
  namespace: monitoring
spec:
  selector:
    app: prometheus
  ports:
    - name: http
      port: 9090
      targetPort: http
  type: ClusterIP

Configuring Applications to Push OTLP

With this deployment, any application in the cluster can push OTLP metrics by setting the following environment variables (works with any OTel SDK supporting OTLP HTTP):

env:
  - name: OTEL_SERVICE_NAME
    valueFrom:
      fieldRef:
        fieldPath: metadata.labels['app']
  - name: OTEL_SERVICE_NAMESPACE
    valueFrom:
      fieldRef:
        fieldPath: metadata.namespace
  - name: OTEL_METRICS_EXPORTER
    value: "otlp"
  - name: OTEL_EXPORTER_OTLP_METRICS_ENDPOINT
    value: "http://prometheus.monitoring.svc.cluster.local:9090/api/v1/otlp/v1/metrics"
  - name: OTEL_EXPORTER_OTLP_METRICS_PROTOCOL
    value: "http/protobuf"
  - name: OTEL_METRIC_EXPORT_INTERVAL
    value: "30000"
  - name: OTEL_RESOURCE_ATTRIBUTES
    value: "deployment.environment=production,k8s.namespace.name=$(NAMESPACE)"

This approach works particularly well in environments using the OTel Operator for Kubernetes, where the Instrumentation CRD can inject these environment variables automatically into pods based on namespace or pod label selectors — zero-touch instrumentation with native Prometheus storage.

Frequently Asked Questions

Can I use Prometheus 3.0 OTLP ingestion for traces and logs?

No. The Prometheus 3.0 OTLP receiver handles only metrics. Prometheus is a metrics store — it has no data model for traces or logs. For traces, you need a backend like Jaeger or Grafana Tempo. For logs, you need Loki, Elasticsearch, or a similar system. The OTel Collector is the appropriate routing layer when you need to send all three signals to their respective backends from a single application-side push endpoint.

Does the kube-prometheus-stack Helm chart support Prometheus 3.0?

Yes, with caveats. The kube-prometheus-stack chart updated its Prometheus image to 3.0 starting with chart version 66.0.0. However, some bundled recording rules and alerting rules may need adjustment for the PromQL changes and default behavioral differences. The Prometheus Operator itself (version 0.78+) has been updated to support the new configuration options including the otlp configuration block. If you are managing Prometheus via the Operator, you will configure OTLP settings through the Prometheus CRD’s spec.additionalArgs and a custom PrometheusConfiguration resource.

What happens to existing Prometheus 2.x metric names when I enable UTF-8 support?

Existing metrics with underscore-based names continue to work exactly as before. Enabling UTF-8 support is purely additive — it allows the storage and querying of metric names containing dots and other UTF-8 characters, but it does not rename or modify existing metrics. Your existing dashboards, alerting rules, and recording rules continue to function without modification. Only metrics ingested via OTLP (or exposed by exporters using OTel naming conventions) will use dot-separated names.

How does native OTLP ingestion affect Prometheus’s pull model?

It coexists with it. Prometheus 3.0 continues to scrape targets via the pull model on the same 9090 port. The OTLP endpoint is an additional ingestion path, not a replacement for scraping. You can have a Prometheus instance simultaneously scraping Kubernetes pods via service discovery and receiving OTLP push metrics from applications — both are stored in the same TSDB and queryable via the same PromQL interface. This hybrid approach is common during migrations, where legacy components are scraped and new OTel-instrumented services push via OTLP.

Is the Prometheus 3.0 OTLP receiver suitable for high-volume production workloads?

For moderate volumes, yes. The OTLP receiver is synchronous — the HTTP request completes only after the samples are written to the WAL. Under very high ingestion rates (hundreds of thousands of samples per second), this can create back-pressure that affects application latency. The OTel Collector handles this better through internal buffering, retry queues, and batch processing. For high-volume scenarios, the recommended pattern is: applications push to OTel Collector (which acknowledges immediately and buffers), Collector pushes to Prometheus via OTLP or remote_write in optimized batches. For the majority of Kubernetes workloads — dozens to hundreds of services with typical metric cardinality — the native OTLP receiver performs well without an intermediary.

Frequently Asked Questions

Can Prometheus 3 receive OpenTelemetry metrics directly?

Yes. Prometheus 3 ships a native OTLP endpoint (/api/v1/otlp/v1/metrics) — an OpenTelemetry Collector or SDK can push metrics straight to it with the OTLP/HTTP exporter, no Prometheus remote-write translation in between.

Does Prometheus 3 support dots in metric names?

Yes — UTF-8 support means OpenTelemetry-style names like http.server.duration no longer have to be flattened to underscores. Querying them needs the quoted syntax: {"http.server.duration"}. Legacy normalization remains available for compatibility with existing dashboards.

Should I push OTLP to Prometheus or scrape with the Collector?

Pushing OTLP fits when OpenTelemetry SDKs are already your instrumentation layer and you want Prometheus as the storage/query engine. Keep native scraping when you depend on up/absent semantics and service discovery — the pull model is still what most alerting assumes. Mixed setups are normal: scrape infrastructure, push application OTLP.

Do existing dashboards break after moving to OTLP ingestion?

Metric names change unless you enable normalization (dots instead of underscores, no automatic _total suffixes), so PromQL that matches on names needs updating. The data itself is fine — it is a naming migration, not a data migration.

Helm Values JSON Schema: Validate Your values.yaml Before It Breaks Production

Helm Values JSON Schema: Validate Your values.yaml Before It Breaks Production

Helm is the de facto package manager for Kubernetes, and values.yaml is its primary interface for configuration. Yet for years, that interface has been completely unvalidated by default — a free-form YAML file where any key can be anything, where typos silently pass through, and where misconfigured deployments only reveal themselves when pods fail to start in production. The values.schema.json file changes that equation entirely. This article explains why schema validation matters, how to implement it properly, and how to integrate it into a modern CI/CD pipeline.

The Problem: Silent Failures in Production

Consider a platform team managing dozens of Helm releases across multiple clusters. A developer submits a values override file with replicaCount: "3" instead of replicaCount: 3 — a string where an integer is expected. Or they set image.pullPolicy: Allways with a typo. Or they omit a required secret reference that the application needs to boot. In all three cases, Helm without schema validation will happily render the templates, produce Kubernetes manifests, and apply them to the cluster. The failure surfaces later — sometimes much later — as a CrashLoopBackOff, an ImagePullBackOff, or a subtle runtime error that takes hours to debug.

This is not a hypothetical scenario. It is the daily reality for teams operating at scale without values validation. The root cause is architectural: Helm templates use Go’s text/template engine, which is weakly typed and permissive by design. A template that does {{ .Values.replicaCount }} will render whether the value is an integer, a string, or even a boolean. The resulting Kubernetes manifest may be invalid, but that error only surfaces when the Kubernetes API server rejects it — or worse, accepts it but interprets it differently than intended.

The consequences compound at scale. When a chart is used by multiple teams, the lack of a formal contract for acceptable values means every consumer has to read through template files and comments to understand what inputs are valid. There is no machine-readable specification. There is no IDE support. There is no guardrail. The only documentation is whatever the chart author happened to write in comments inside values.yaml — and comments do not stop a CI pipeline from shipping a broken deployment.

What Is values.schema.json

Since Helm 3.0.0, released in November 2019, Helm supports an optional values.schema.json file at the root of a chart directory — the same level as Chart.yaml and values.yaml. This file is a JSON Schema draft-07 document that formally describes the structure, types, constraints, and required fields for the chart’s values.

When this file is present, Helm automatically validates the merged values (defaults from values.yaml merged with any user-supplied overrides) against the schema at multiple points: during helm install, helm upgrade, helm template, and helm lint. If validation fails, Helm refuses to proceed and prints a human-readable error message identifying exactly which value failed and why. This transforms a class of runtime failures into build-time failures — the correct direction for any production system.

The choice of JSON Schema draft-07 specifically is worth noting. Draft-07 is widely supported by tooling, including the Red Hat YAML extension for VS Code, JetBrains IDEs, and most JSON Schema validators. It introduced the if/then/else conditional keywords that are particularly useful for Helm charts. More recent drafts (2019-09, 2020-12) offer additional features but have less universal tooling support, making draft-07 the pragmatic choice for chart authors today.

Chart Directory Structure

my-app/
├── Chart.yaml
├── values.yaml
├── values.schema.json      ← lives here
├── charts/
└── templates/
    ├── deployment.yaml
    ├── service.yaml
    ├── ingress.yaml
    └── _helpers.tpl

The schema file is included when a chart is packaged with helm package and distributed through chart repositories. Consumers of the chart get schema validation automatically without any additional configuration — the guardrails ship with the chart itself.

How Helm Uses the Schema

Helm’s validation behavior is straightforward but has some nuances worth understanding. When Helm processes a release, it first merges all value sources in order of increasing precedence: chart defaults (values.yaml), parent chart values, -f value files, and finally --set flags. The merged result is then validated against the schema as a single operation.

This means the schema validates the effective values, not each source in isolation. A required field that has a default in values.yaml will pass validation even when not specified by the user, because the merged result includes the default. This is the correct behavior — it validates what will actually be used during rendering.

The validation happens before template rendering. If schema validation fails, Helm exits with a non-zero status code and prints all validation errors. The error output is structured and actionable:

$ helm install my-release ./my-app --set replicaCount=abc

Error: values don't meet the specifications of the schema(s) in the following chart(s):
my-app:
- replicaCount: Invalid type. Expected: integer, given: string

For helm lint, which is typically used in CI pipelines without installing to a cluster, schema validation also runs. This makes helm lint a powerful pre-deployment gate when schema files are present.

IDE Benefits: Autocompletion and Inline Validation

Beyond Helm’s own validation, values.schema.json unlocks IDE support that significantly improves the developer experience when working with values files. The Red Hat YAML extension for VS Code can reference a JSON Schema file to provide autocompletion, type checking, and inline error highlighting for YAML files.

To enable this, add a yaml.schemas configuration to your VS Code workspace settings or the user settings file:

// .vscode/settings.json
{
  "yaml.schemas": {
    "./my-app/values.schema.json": "./my-app/values.yaml"
  }
}

With this configuration, editing values.yaml in VS Code will show autocompletion for defined keys, inline errors for type mismatches, and hover documentation pulled from the description fields in your schema. For platform teams maintaining internal Helm charts, this transforms the chart into a self-documenting, IDE-aware configuration interface — without any additional tooling investment.

JetBrains IDEs (IntelliJ IDEA, GoLand, etc.) support JSON Schema associations through the Languages & Frameworks > Schemas and DTDs > JSON Schema Mappings settings panel, providing equivalent functionality for teams using those tools.

Building the Schema: A Practical Guide

Let’s build a complete, realistic example. Start with a typical values.yaml for a web application chart:

# values.yaml
replicaCount: 2

image:
  repository: myorg/my-app
  tag: "1.0.0"
  pullPolicy: IfNotPresent

service:
  type: ClusterIP
  port: 80

ingress:
  enabled: false
  hostname: ""
  tls: false

resources:
  requests:
    cpu: "100m"
    memory: "128Mi"
  limits:
    cpu: "500m"
    memory: "512Mi"

autoscaling:
  enabled: false
  minReplicas: 1
  maxReplicas: 10
  targetCPUUtilizationPercentage: 80

config:
  logLevel: info
  databaseUrl: ""

nodeSelector: {}
tolerations: []
affinity: {}

Now the full values.schema.json that validates this structure:

{
  "$schema": "http://json-schema.org/draft-07/schema#",
  "title": "my-app Helm Chart Values",
  "description": "Configuration values for the my-app Helm chart",
  "type": "object",
  "additionalProperties": false,
  "required": ["image", "service"],
  "$defs": {
    "resourceQuantity": {
      "type": "string",
      "pattern": "^[0-9]+(\\.[0-9]+)?(m|Ki|Mi|Gi|Ti|Pi|Ei|k|M|G|T|P|E)?$",
      "description": "A Kubernetes resource quantity (e.g. 100m, 128Mi, 1Gi)"
    }
  },
  "properties": {
    "replicaCount": {
      "type": "integer",
      "minimum": 0,
      "maximum": 50,
      "default": 2,
      "description": "Number of pod replicas. Set to 0 to scale down."
    },
    "image": {
      "type": "object",
      "additionalProperties": false,
      "required": ["repository", "tag"],
      "description": "Container image configuration",
      "properties": {
        "repository": {
          "type": "string",
          "minLength": 1,
          "description": "Container image repository"
        },
        "tag": {
          "type": "string",
          "pattern": "^[a-zA-Z0-9._-]+$",
          "minLength": 1,
          "description": "Image tag. Avoid using 'latest' in production."
        },
        "pullPolicy": {
          "type": "string",
          "enum": ["Always", "IfNotPresent", "Never"],
          "default": "IfNotPresent",
          "description": "Kubernetes imagePullPolicy"
        }
      }
    },
    "service": {
      "type": "object",
      "additionalProperties": false,
      "required": ["type", "port"],
      "properties": {
        "type": {
          "type": "string",
          "enum": ["ClusterIP", "NodePort", "LoadBalancer", "ExternalName"],
          "description": "Kubernetes Service type"
        },
        "port": {
          "type": "integer",
          "minimum": 1,
          "maximum": 65535,
          "description": "Service port"
        }
      }
    },
    "ingress": {
      "type": "object",
      "additionalProperties": false,
      "properties": {
        "enabled": {
          "type": "boolean",
          "default": false
        },
        "hostname": {
          "type": "string",
          "description": "Ingress hostname. Required when ingress.enabled is true."
        },
        "tls": {
          "type": "boolean",
          "default": false,
          "description": "Enable TLS for the ingress"
        }
      },
      "if": {
        "properties": {
          "enabled": { "const": true }
        },
        "required": ["enabled"]
      },
      "then": {
        "required": ["hostname"],
        "properties": {
          "hostname": {
            "minLength": 1,
            "pattern": "^[a-zA-Z0-9]([a-zA-Z0-9\\-\\.]+)?[a-zA-Z0-9]$"
          }
        }
      }
    },
    "resources": {
      "type": "object",
      "additionalProperties": false,
      "properties": {
        "requests": {
          "type": "object",
          "additionalProperties": false,
          "properties": {
            "cpu": { "$ref": "#/$defs/resourceQuantity" },
            "memory": { "$ref": "#/$defs/resourceQuantity" }
          }
        },
        "limits": {
          "type": "object",
          "additionalProperties": false,
          "properties": {
            "cpu": { "$ref": "#/$defs/resourceQuantity" },
            "memory": { "$ref": "#/$defs/resourceQuantity" }
          }
        }
      }
    },
    "autoscaling": {
      "type": "object",
      "additionalProperties": false,
      "properties": {
        "enabled": {
          "type": "boolean",
          "default": false
        },
        "minReplicas": {
          "type": "integer",
          "minimum": 1
        },
        "maxReplicas": {
          "type": "integer",
          "minimum": 1,
          "maximum": 100
        },
        "targetCPUUtilizationPercentage": {
          "type": "integer",
          "minimum": 1,
          "maximum": 100
        }
      },
      "if": {
        "properties": {
          "enabled": { "const": true }
        },
        "required": ["enabled"]
      },
      "then": {
        "required": ["minReplicas", "maxReplicas"]
      }
    },
    "config": {
      "type": "object",
      "additionalProperties": false,
      "properties": {
        "logLevel": {
          "type": "string",
          "enum": ["debug", "info", "warn", "error"],
          "default": "info",
          "description": "Application log level"
        },
        "databaseUrl": {
          "type": "string",
          "description": "Database connection URL"
        }
      }
    },
    "nodeSelector": {
      "type": "object",
      "description": "Node selector labels for pod scheduling"
    },
    "tolerations": {
      "type": "array",
      "description": "Pod tolerations"
    },
    "affinity": {
      "type": "object",
      "description": "Pod affinity rules"
    }
  }
}

Key Schema Patterns Explained

additionalProperties: false

This is arguably the most important pattern in a Helm schema. Without it, unknown keys pass validation silently — which defeats much of the purpose. With "additionalProperties": false, any key not listed in properties causes a validation error. This catches typos like repicaCount instead of replicaCount, which would otherwise silently use the default value and leave the developer wondering why their override had no effect.

Apply it at every nested object level, not just the root. A typo inside image: or resources: is just as dangerous as one at the top level.

$defs for Reusable Definitions

The $defs keyword (called definitions in earlier draft versions, though draft-07 supports both) provides a namespace for reusable schema fragments. In the example above, resourceQuantity is defined once and referenced via $ref in both requests and limits. This avoids duplication and ensures consistent validation logic across related fields.

For larger charts, $defs becomes essential. Common patterns include reusable schemas for image configurations, resource requirements, probe configurations, and environment variable maps.

Conditional Validation with if/then/else

The if/then/else construct in JSON Schema draft-07 is particularly powerful for Helm charts, where many values are conditional on a feature toggle. The ingress example above demonstrates this: when ingress.enabled is true, the hostname field becomes required and must match a valid hostname pattern. When ingress is disabled, the hostname can be empty or omitted entirely.

This pattern can be extended for more complex scenarios. For example, enforcing that when autoscaling.enabled is true, the standalone replicaCount should not be set (since the HPA controls replica count):

{
  "if": {
    "properties": {
      "autoscaling": {
        "properties": {
          "enabled": { "const": true }
        },
        "required": ["enabled"]
      }
    }
  },
  "then": {
    "properties": {
      "replicaCount": {
        "description": "replicaCount is ignored when autoscaling is enabled"
      }
    }
  }
}

Pattern Validation for Image Tags

The image tag field is a common source of production issues. Teams accidentally deploy with latest, which is non-deterministic and makes rollbacks unreliable. A pattern constraint can enforce semantic versioning or at least ban the latest tag in production charts:

"tag": {
  "type": "string",
  "not": {
    "enum": ["latest", ""]
  },
  "pattern": "^[0-9]+\\.[0-9]+\\.[0-9]+",
  "description": "Semantic version tag required. 'latest' is not permitted."
}

This enforces that image tags start with a semantic version number, immediately rejecting latest, empty strings, or arbitrary branch names that would produce non-reproducible deployments.

Enum for Controlled Vocabularies

Fields with a fixed set of valid values — Kubernetes service types, image pull policies, log levels — should use enum. This is more precise than a pattern and produces clearer error messages. It also enables IDE autocompletion to show exactly the valid options as a pick-list, rather than requiring the developer to remember or look up acceptable values.

CI/CD Integration

GitHub Actions

The most direct integration point is helm lint, which runs schema validation as part of its checks. A minimal GitHub Actions workflow that validates a chart on every pull request looks like this:

# .github/workflows/helm-lint.yaml
name: Helm Lint

on:
  pull_request:
    paths:
      - 'charts/**'

jobs:
  lint:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4

      - name: Set up Helm
        uses: azure/setup-helm@v4
        with:
          version: '3.14.0'

      - name: Lint chart with default values
        run: helm lint charts/my-app

      - name: Lint chart with staging values
        run: helm lint charts/my-app -f charts/my-app/ci/staging-values.yaml

      - name: Lint chart with production values
        run: helm lint charts/my-app -f charts/my-app/ci/production-values.yaml

      - name: Validate template rendering
        run: |
          helm template my-app charts/my-app \
            -f charts/my-app/ci/production-values.yaml \
            --debug > /dev/null

The ci/ directory convention (values files specifically for CI testing) is a pattern from the chart-testing tool and works well for validating multiple realistic value combinations, not just the defaults.

For teams using the ct (chart-testing) CLI tool from the Helm project, schema validation is automatically included in the ct lint command, which also handles chart versioning checks and YAML linting:

      - name: Chart Testing lint
        uses: helm/chart-testing-action@v2.6.1

      - name: Run chart-testing lint
        run: ct lint --target-branch ${{ github.event.repository.default_branch }}

Pre-commit Hooks

For local development, pre-commit hooks catch issues before code is even pushed. The pre-commit framework makes this straightforward:

# .pre-commit-config.yaml
repos:
  - repo: https://github.com/gruntwork-io/pre-commit
    rev: v0.1.23
    hooks:
      - id: helmlint

  - repo: local
    hooks:
      - id: helm-schema-validate
        name: Helm Schema Validation
        language: script
        entry: scripts/validate-helm-schemas.sh
        files: ^charts/.*values.*\.yaml$
#!/usr/bin/env bash
# scripts/validate-helm-schemas.sh
set -euo pipefail

for chart_dir in charts/*/; do
  if [[ -f "${chart_dir}/values.schema.json" ]]; then
    echo "Linting ${chart_dir}..."
    helm lint "${chart_dir}" --strict
  fi
done

ArgoCD and Flux Integration

Both ArgoCD and Flux (Helm Controller) invoke helm template internally when reconciling Helm releases. Since helm template runs schema validation when a schema file is present, any invalid values in an HelmRelease or ArgoCD Application manifest will cause the reconciliation to fail with a clear error message — visible in the controller logs and surface as a degraded resource status. No additional configuration is required; schema validation is automatic.

Generating Schemas from Existing Charts

For charts that already have a well-structured values.yaml, writing a schema from scratch is time-consuming but not starting from zero. Several tools can generate a draft schema that you then refine:

  • helm-values-schema-json — a Helm plugin (helm plugin install https://github.com/losisin/helm-values-schema-json) that introspects values.yaml and generates a draft schema with inferred types. Run with helm schema-gen values.yaml.
  • json-schema-generator online tools — paste your values as JSON (convert YAML to JSON first) and get a draft schema back.
  • Manually from scratch — for new charts, writing the schema alongside the values file from the beginning is the most accurate approach and requires no extra tooling.

Generated schemas are always starting points. They infer types from existing values but cannot know about intended constraints, enums, patterns, required fields in conditional cases, or additionalProperties: false at nested levels. Manual review and refinement is always necessary.

Common Mistakes and How to Avoid Them

MistakeSymptomFix
Missing additionalProperties: falseTypos in key names pass validation silentlyAdd it at every object level, including nested objects
Schema only at root levelNested typos go undetectedApply additionalProperties: false recursively
Not including defaults in schemaIDE shows fields as required when they are optionalAdd default to all optional fields
Overly strict patterns blocking valid valuesLegitimate deployments fail schema validationTest patterns against your real value space before shipping
Using definitions instead of $defsWorks in most tools but is draft-2019-09+ terminologyUse $defs for draft-07 compliance; both work in practice
Schema not committed to the chart repoConsumers get no validation when pulling from repositoryAlways commit values.schema.json alongside the chart
Validating subchart values through parent schemaSchema errors for subchart values the parent doesn’t ownDo not attempt to validate subchart values in parent schema; each chart owns its own schema

The Null Value Problem

A subtle but common issue: in YAML, an unset key with no value (key:) resolves to null, not an empty string or zero. If your schema defines a field as "type": "string", a null value will fail validation. To handle optional fields that users might leave blank, use a type union:

"databaseUrl": {
  "type": ["string", "null"],
  "description": "Database connection URL. Leave null to use the default."
}

Alternatively, ensure your values.yaml defaults use empty strings ("") rather than bare keys, and document that convention for chart consumers.

Schema Drift

As charts evolve, new values get added to values.yaml without corresponding updates to values.schema.json. Over time the schema becomes stale and provides partial coverage. The fix is procedural: treat schema updates as part of the definition of done for any PR that modifies values. Code review should include checking that new or modified values have corresponding schema entries.

Frequently Asked Questions

Does values.schema.json validate subchart values?

No. Each chart in a dependency relationship validates only its own values against its own schema. If chart A depends on chart B, and chart B has a schema, chart B’s schema validates the values under the b: key in chart A’s values.yaml — but only when processed in the context of chart B itself. Chart A’s schema should not attempt to describe chart B’s values structure. This is by design: it maintains loose coupling between charts and allows subcharts to evolve their schemas independently.

Can I use JSON Schema draft-2020-12 instead of draft-07?

Technically, Helm does not strictly enforce which draft version you use — it uses the Go library github.com/xeipuuv/gojsonschema, which supports draft-04 through draft-07. Using newer draft keywords that are not supported by this library may cause them to be silently ignored rather than throwing an error. For IDE support, draft-07 has the broadest compatibility. If you need features from newer drafts (like unevaluatedProperties from 2020-12), test carefully to confirm they are enforced by Helm’s validator and not silently skipped.

How do I handle values that differ between environments without schema conflicts?

The schema should describe all valid values across all environments. Use enum to enumerate all valid values for a field, and use if/then/else for constraints that only apply in certain configurations. The schema is a contract for what the chart accepts, not a policy for what a specific environment should use. Environment-specific policies (such as “production must use a minimum of 3 replicas”) are better enforced at a higher level — through admission controllers like OPA Gatekeeper or Kyverno — rather than in the chart schema itself.

Does schema validation run when using helm template for dry runs?

Yes. helm template runs schema validation before rendering templates. This makes it useful as a validation step in CI pipelines even without a live cluster: helm template release-name ./chart -f values-override.yaml will fail with schema errors if the values are invalid, and will output the rendered manifests if they are valid. Piping the output to kubectl apply --dry-run=client -f - adds an additional layer of Kubernetes API validation for a thorough offline check.

Should I add values.schema.json to charts I don’t maintain (upstream charts)?

For upstream charts you consume but do not maintain (such as Bitnami charts, ingress-nginx, cert-manager), the recommended approach is to maintain a separate JSON Schema file in your own GitOps repository that validates your specific values overlay files. Tools like jsonschema (Python) or ajv (Node.js) can validate a YAML/JSON values file against a schema in CI without Helm being involved. This gives you schema validation for your environment-specific overrides without needing to modify upstream chart sources.

Frequently Asked Questions

What is values.schema.json in Helm?

A JSON Schema file placed next to values.yaml in a chart. Helm validates the merged values against it on helm install, upgrade, lint and template, and refuses to render if validation fails — turning typos and wrong types into hard errors before anything reaches the cluster.

Does Helm validate values.schema.json automatically?

Yes. If the file exists in the chart root, validation runs on every install, upgrade, lint and template with no extra flags. It validates the final merged values — defaults plus every -f file plus --set overrides.

How do I reject unknown keys in Helm values?

Set "additionalProperties": false at the levels you want closed. With that, a misspelled key like replicaCont fails validation instead of being silently ignored — the single most valuable line in most schemas.

Do subcharts have their own values schema?

Yes, and each subchart’s schema validates its own slice of the values. A parent chart cannot loosen a subchart’s schema; if a dependency ships a strict schema, your overrides must comply with it.

After NGINX Ingress Controller: Alternatives and Migration Guide

After NGINX Ingress Controller: Alternatives and Migration Guide

If you manage Kubernetes clusters in production, the last 18 months have been uncomfortable. Two of the most widely deployed NGINX-based Ingress Controllers have faced critical security vulnerabilities, deprecation announcements, and shifting maintenance responsibilities — all while the Kubernetes project accelerates its push toward a new traffic management standard. This is not a drill. Teams running ingress-nginx or the F5/NGINX Ingress Controller need a clear picture of what changed, what it means for their clusters, and what their realistic options are going forward.

First, Clear the Confusion: There Are Two NGINX Ingress Controllers

One of the most persistent sources of confusion in the Kubernetes networking space is that there are two completely different projects both called “NGINX Ingress Controller,” maintained by different organizations, with different architectures and different licensing.

ingress-nginx (kubernetes/ingress-nginx)

This is the community-maintained controller under the Kubernetes project umbrella, hosted at github.com/kubernetes/ingress-nginx. It uses the open-source NGINX as its data plane, configured via Lua scripting and dynamically generated nginx.conf files. This is the controller most teams end up with when they follow the official Kubernetes documentation or install from the Helm chart referenced in the ingress guide. It is free, open-source, and until recently was considered the default choice.

NGINX Ingress Controller (nginxinc/kubernetes-ingress)

This is the commercial and open-source controller maintained by F5/NGINX, hosted at github.com/nginxinc/kubernetes-ingress. It also supports NGINX Plus (the commercial version with enhanced features like active health checks, JWT authentication, and advanced load balancing). The architecture is different — it uses native NGINX APIs rather than the Lua-heavy approach — and it targets enterprise customers looking for support contracts and advanced capabilities.

These two controllers are not interchangeable. Configuration annotations differ, Helm chart values differ, and behavior under edge cases differs substantially. Understanding which one your cluster runs is the necessary starting point for any decision about migration.

# Check which NGINX IC you are actually running
kubectl get pods -n ingress-nginx -o jsonpath='{.items[*].spec.containers[*].image}'

# Community controller image looks like:
# registry.k8s.io/ingress-nginx/controller:v1.x.x

# F5/NGINX controller image looks like:
# nginx/nginx-ingress:x.x.x  or  private-registry.nginx.com/nginx-ic/nginx-plus-ingress:x.x.x

What Actually Happened: A Timeline of Disruption

The ingress-nginx CVEs (2024)

In March 2024, security researchers disclosed a set of critical vulnerabilities in ingress-nginx under the collective name IngressNightmare (CVE-2025-1097, CVE-2025-1098, CVE-2025-1974, CVE-2025-24514). The most severe of these, rated CVSS 9.8, allowed unauthenticated remote code execution against the ingress-nginx admission webhook. An attacker with network access to the admission controller could craft a malicious Ingress object to inject arbitrary NGINX configuration, ultimately achieving code execution in the controller pod — which in many clusters runs with elevated permissions and access to service account tokens across namespaces.

The vulnerabilities affected the vast majority of ingress-nginx deployments in the wild. Wiz Research, which discovered and disclosed the issues, estimated that approximately 43% of cloud environments were exposed. Patches were released in versions 1.11.5 and 1.12.1, but the incident forced uncomfortable questions about the controller’s security posture and the architecture decisions (particularly the admission webhook design) that made it possible.

Maintenance Concerns in ingress-nginx

Beyond the CVEs, the ingress-nginx project has faced ongoing concerns about maintainer bandwidth. The project is maintained by a small group of volunteers and relies heavily on community contributions. Issue response times slowed, pull requests aged, and the pace of feature development declined relative to alternatives. For a component as critical as the cluster ingress layer, this created legitimate concern about long-term sustainability without corporate backing or broader contributor growth.

F5/NGINX Deprecation Announcement

On the commercial side, F5/NGINX announced in early 2025 that the nginxinc/kubernetes-ingress controller — particularly its open-source tier — would undergo significant changes. F5 signaled a strategic shift toward NGINX Gateway Fabric, their implementation of the Kubernetes Gateway API specification. The message was clear: investment in the Ingress-based controller would be reduced, and customers were encouraged to plan migrations toward Gateway API-native solutions.

For teams running NGINX Plus-based ingress with support contracts, this was a significant business concern. The product they had licensed and standardized on was being steered toward end-of-life on the Ingress API, even if exact timelines remained somewhat ambiguous in the initial announcements.

Real Impact on Production Clusters

The practical consequences depend heavily on which controller you run and how your clusters are configured. Here is an honest assessment:

Immediate Security Risk

If you run ingress-nginx and have not patched to 1.11.5+ or 1.12.1+, your admission webhook is a critical attack surface. Patching is non-negotiable and should have happened already. The admission webhook can be disabled if you are not using it for validation (many teams are not), which significantly reduces the attack surface while you plan a longer-term migration.

# Check your current ingress-nginx version
kubectl get deployment ingress-nginx-controller -n ingress-nginx \
  -o jsonpath='{.spec.template.spec.containers[0].image}'

# Verify admission webhook is configured
kubectl get validatingwebhookconfigurations | grep ingress

# If you need to disable the webhook temporarily (reduces but does not eliminate risk):
kubectl delete validatingwebhookconfiguration ingress-nginx-admission

Operational Uncertainty

Even after patching, the underlying questions remain. Teams are now asking: should we invest in hardening and tuning ingress-nginx knowing it may not be the strategic direction? Should we migrate now when it is our choice, rather than later when it may be forced? For NGINX IC customers, they are evaluating whether their licensing costs justify continued investment in a product being steered toward deprecation.

Configuration Migration Complexity

The real cost of migration is in the annotation-heavy configurations that accumulate over time. Teams that have built complex routing logic using nginx.ingress.kubernetes.io/* annotations — custom headers, rate limiting, auth snippets, rewrite rules, canary traffic splitting — face significant rework when switching controllers. This is the primary reason many teams are reluctant to move despite clear signals that a transition is coming.

The Alternatives: An Honest Evaluation

There is no shortage of Ingress controller options. The question is which alternatives are mature enough for production workloads at scale, and what trade-offs each brings.

Traefik

Traefik Proxy (and its Kubernetes-native version via Traefik Hub) has emerged as the most popular alternative for teams leaving ingress-nginx. It supports the standard Kubernetes Ingress API for drop-in compatibility, its own IngressRoute CRDs for advanced features, and Kubernetes Gateway API. It is written in Go, has strong TLS automation via Let’s Encrypt, and has excellent observability with built-in metrics and a real-time dashboard.

Trade-offs: Traefik’s configuration model is different enough from NGINX that complex routing logic requires rethinking rather than translating. Performance under very high connection counts is generally good but NGINX has a longer track record in extreme-scale deployments. The commercial offering (Traefik Hub) adds API gateway capabilities but introduces vendor dependency.

Envoy Gateway

Envoy Gateway is now a CNCF project and implements the Kubernetes Gateway API natively using Envoy as its data plane. This is arguably the most strategically aligned option for teams that want to bet on the future of Kubernetes networking. Envoy is battle-tested (it powers Istio, Contour, and large-scale service meshes at companies like Lyft and Google), and the Gateway API implementation is comprehensive and actively developed.

Trade-offs: Envoy Gateway is relatively young as a standalone project. Teams unfamiliar with Envoy will face a steeper learning curve for debugging and custom configuration. The operational model differs significantly from NGINX-based controllers. However, for greenfield deployments or teams willing to invest in the transition, this is a strong forward-looking choice.

Cilium Gateway API

If your cluster already runs Cilium as the CNI, enabling Gateway API support is a natural evolution. Cilium’s Gateway API implementation leverages eBPF for high-performance packet processing, avoiding the overhead of userspace proxy hops entirely. It is deeply integrated with Cilium’s network policy model and observability stack (Hubble).

Trade-offs: This option is only relevant if you are already committed to Cilium as your CNI, or are willing to make that switch simultaneously. Migrating both the CNI and the ingress layer at the same time is a significant operational risk. For Cilium shops, however, this consolidates complexity and provides excellent performance and observability.

HAProxy Ingress

HAProxy Ingress Controller is maintained by the HAProxy Technologies team and has a strong reputation for raw performance and precise traffic control. It supports both Ingress and Gateway API and has a long track record in high-throughput production environments. For teams with existing HAProxy expertise, it provides a familiar mental model for load balancing configuration.

Trade-offs: Smaller community than Traefik or NGINX. Less ecosystem tooling and fewer tutorials. Best suited for teams that specifically want HAProxy’s capabilities (fine-grained connection management, advanced health checking, TCP/HTTP mode flexibility) rather than as a default choice.

Kong Ingress Controller

Kong bridges the gap between an Ingress controller and a full API gateway. It supports Ingress and Gateway API resources alongside its own Kong-native plugin system for authentication, rate limiting, transformation, and observability. For teams that need API gateway capabilities rather than pure L7 routing, Kong provides a unified platform.

Trade-offs: Kong adds operational complexity. Running Kong requires either a PostgreSQL database (DB-mode) or careful management of declarative configuration (DB-less mode). The plugin ecosystem is powerful but introduces additional configuration surface. For teams that just need ingress routing, Kong may be more than necessary. For teams building API platforms, it is worth the overhead.

Istio Gateway

Istio’s ingress gateway (now aligned with Gateway API via its Kubernetes Gateway API integration) provides entry-point traffic management as part of a full service mesh. If your organization is planning or running Istio for east-west traffic, using Istio’s gateway for north-south traffic creates a unified data plane (Envoy) and consistent observability across all service communication.

Trade-offs: Istio is a serious operational commitment. The control plane overhead, the learning curve, and the impact on pod scheduling and sidecar management are significant. Choosing Istio purely for ingress replacement is like buying a race car because you needed a vehicle with good brakes. Consider this path only if service mesh capabilities are on your roadmap.

NGINX Gateway Fabric (F5’s Gateway API implementation)

F5/NGINX is building NGINX Gateway Fabric as their strategic forward path — an NGINX-based implementation of the Kubernetes Gateway API. For teams heavily invested in NGINX and wanting to stay in that ecosystem while moving to Gateway API, this provides a migration path within familiar territory. It is still maturing but represents where F5 is putting its development resources.

Comparison Matrix

ControllerIngress APIGateway APIMaturityBest ForComplexity
ingress-nginxYesPartialHighExisting deployments, familiar configLow
TraefikYesYesHighGeneral purpose, rapid migrationLow-Medium
Envoy GatewayNoYes (native)MediumGreenfield, future-alignedMedium
Cilium GatewayYesYesMediumCilium CNI clustersLow (if Cilium)
HAProxy IngressYesYesHighHigh-throughput, HAProxy expertiseMedium
KongYesYesHighAPI gateway requirementsHigh
Istio GatewayVia Gateway APIYesHighService mesh adoptersVery High
NGINX Gateway FabricNoYes (native)Low-MediumNGINX shops moving to Gateway APIMedium

Gateway API: The Strategic Direction You Cannot Ignore

The Kubernetes Gateway API is not simply “Ingress v2.” It is a fundamentally richer traffic management model designed to address the limitations that drove teams to annotation-based workarounds for the past several years. Understanding it is essential regardless of which controller you choose, because the ecosystem is clearly converging on it.

The core resource hierarchy consists of GatewayClass (defines a type of gateway, created by infrastructure providers), Gateway (a specific instance of a listener configuration, typically managed by platform teams), and HTTPRoute, TCPRoute, GRPCRoute, and other route resources (managed by application teams). This separation of concerns maps cleanly onto organizational roles — infrastructure teams control the gateway, application teams control their routing rules.

# Example Gateway API resources replacing an ingress-nginx Ingress
apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata:
  name: production-gateway
  namespace: infra
spec:
  gatewayClassName: nginx  # or envoy, traefik, cilium, etc.
  listeners:
  - name: https
    protocol: HTTPS
    port: 443
    tls:
      mode: Terminate
      certificateRefs:
      - name: wildcard-tls
    allowedRoutes:
      namespaces:
        from: All
---
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
  name: my-app
  namespace: my-app-namespace
spec:
  parentRefs:
  - name: production-gateway
    namespace: infra
  hostnames:
  - "app.example.com"
  rules:
  - matches:
    - path:
        type: PathPrefix
        value: /api
    backendRefs:
    - name: my-api-service
      port: 8080
  - matches:
    - path:
        type: PathPrefix
        value: /
    backendRefs:
    - name: my-frontend-service
      port: 3000

Gateway API reached v1.0 (GA for HTTPRoute and Gateway) in October 2023, and v1.1 followed in 2024 with GRPCRoute graduation and expanded features. The project has broad support across controllers (Traefik, Envoy Gateway, Cilium, NGINX Gateway Fabric, Kong, Istio, and others all implement it). The Ingress API is not being removed from Kubernetes, but new feature development is effectively frozen — Gateway API is where capabilities like traffic weighting, header manipulation, request mirroring, and backend protocol configuration are being built.

Decision Framework: Stay, Migrate, or Evaluate?

There is no universal right answer. The following framework helps teams make a context-appropriate decision rather than following hype or panic.

Stay on ingress-nginx if:

  • You have patched to 1.11.5+ or 1.12.1+ and have disabled or hardened the admission webhook
  • Your cluster is stable, heavily annotation-dependent, and migration cost outweighs risk
  • You have internal NGINX expertise and can take ownership of monitoring the project’s maintenance health
  • Your organization has a short-term horizon (decommissioning or major platform change within 12-18 months)

Migrate now if:

  • You are running the F5/NGINX IC with a support contract that is being deprecated
  • Your cluster has moderate annotation complexity and you have engineering cycles available
  • You are planning a major Kubernetes version upgrade or cluster rebuild — do it at the same time
  • Your security team has flagged the CVE history as unacceptable for your risk profile
  • You are building a new cluster or platform team and want to standardize on Gateway API from the start

Evaluate before committing if:

  • Your workloads have complex traffic requirements (WebSockets, gRPC, canary deployments, header-based routing) that differ significantly across controllers
  • You are considering Gateway API but the specific controllers in your environment have not graduated their Gateway API implementations yet
  • You have multi-cluster or multi-tenant requirements that change the analysis
  • You need to assess total cost including commercial support, tooling changes, and team retraining

Migration Checklist

For teams that have decided to migrate, the following sequence reduces risk and ensures nothing critical is missed:

Phase 1: Inventory and Assessment

  • Enumerate all Ingress resources across all namespaces and document their annotations
  • Identify annotations with no direct equivalent in your target controller
  • Map TLS certificate sources (cert-manager, Secrets, external providers) and confirm compatibility
  • Document any custom NGINX configuration snippets (nginx.ingress.kubernetes.io/configuration-snippet, server-snippet) — these are high-risk items that require manual translation
  • Inventory any rate limiting, authentication, or WAF configurations layered on the controller
# Enumerate all ingress resources and their annotations across the cluster
kubectl get ingress -A -o json | jq -r '
  .items[] |
  {
    namespace: .metadata.namespace,
    name: .metadata.name,
    annotations: (.metadata.annotations // {} | keys)
  }
'

Phase 2: Target Controller Validation

  • Deploy target controller to a non-production cluster with identical Ingress/HTTPRoute resources
  • Validate TLS termination, redirect behavior, and timeout configurations
  • Run load tests to confirm performance characteristics match expectations
  • Validate observability — metrics, logs, and traces integrate with your existing stack
  • Test failure scenarios: backend unavailability, certificate expiry, controller pod restart

Phase 3: Staged Production Migration

  • Deploy new controller to production alongside existing controller (different IngressClass)
  • Migrate low-risk, low-traffic Ingress resources first by updating their ingressClassName
  • Use DNS-based canary switching (weighted routing at the DNS level) rather than switching entire IngressClass at once
  • Monitor error rates and latency for 24-48 hours after each batch migration
  • Migrate critical services during low-traffic windows with rollback plan documented
  • Decommission old controller only after all resources are migrated and validated
# Migrate individual Ingress to new controller by changing ingressClassName
kubectl patch ingress my-app -n my-namespace \
  --type='json' \
  -p='[{"op": "replace", "path": "/spec/ingressClassName", "value": "traefik"}]'

# Or if migrating to Gateway API, create equivalent HTTPRoute first,
# test it, then remove the old Ingress resource
kubectl apply -f my-app-httproute.yaml
# Validate, then:
kubectl delete ingress my-app -n my-namespace

Phase 4: Gateway API Adoption (Optional but Recommended)

  • Install Gateway API CRDs if not already present (kubectl apply -f https://github.com/kubernetes-sigs/gateway-api/releases/download/v1.1.0/standard-install.yaml)
  • Define GatewayClass resources matching your chosen controller
  • Migrate Ingress resources to HTTPRoute progressively, starting with simpler configurations
  • Update CI/CD pipelines and Helm charts to generate HTTPRoute instead of Ingress resources for new services
  • Establish a policy: new services use Gateway API; legacy services migrate on their next significant update

Recommendation

For most platform engineering teams reading this in 2025, the pragmatic recommendation is as follows:

Short term (next 30 days): Patch ingress-nginx to the latest release if you are still on it. Assess and harden or disable the admission webhook. This is not optional.

Medium term (3-6 months): Evaluate Traefik or Envoy Gateway against your specific workload requirements. Traefik is the lower-friction migration for teams coming from ingress-nginx on the Ingress API. Envoy Gateway is the stronger strategic choice if you are willing to commit to Gateway API fully. Either way, run a parallel deployment in a non-production environment and measure the delta in operational overhead.

Long term (6-18 months): Plan migration to Gateway API resources regardless of which data plane you choose. The Ingress API will not disappear overnight, but feature parity with Gateway API capabilities will never arrive. Teams that standardize on Gateway API now build the institutional knowledge that will be valuable as the ecosystem continues to evolve.

If you are running F5/NGINX IC under a support contract: engage your F5 account team now to get a clear timeline on the deprecation path and evaluate NGINX Gateway Fabric as a within-ecosystem migration before looking at alternatives. The question is not whether to migrate but when and to what.

Avoid the temptation to treat this as a purely technical decision. The switch of an ingress controller touches CI/CD pipelines, monitoring dashboards, runbooks, on-call playbooks, and engineering team knowledge. Factor in the total transition cost, not just the YAML changes.

Frequently Asked Questions

Is the Kubernetes Ingress API being deprecated or removed?

No. The networking.k8s.io/v1 Ingress API is not deprecated and there are no current plans to remove it from Kubernetes. It will continue to work. What is happening is that the Kubernetes SIG Network has frozen new feature development on the Ingress API and is directing all new traffic management capabilities to Gateway API. In practical terms, if you need a capability that Ingress does not currently provide, you will not get it through Ingress. You will need Gateway API. Existing Ingress resources will continue to function for the foreseeable future.

Can I run two Ingress controllers simultaneously during migration?

Yes, and this is the recommended approach for production migrations. Kubernetes supports multiple IngressClass resources in a cluster, each backed by a different controller. Ingress resources select their controller via the spec.ingressClassName field (or the legacy kubernetes.io/ingress.class annotation). You can run ingress-nginx and Traefik side-by-side, migrating individual Ingress resources by updating their ingressClassName. Once migration is complete and validated, decommission the old controller. Just ensure both controllers are not both marked as the default IngressClass simultaneously, as this causes conflicts.

What happens to cert-manager if I switch controllers?

cert-manager is independent of your Ingress controller and will continue to work regardless of which controller you use. The HTTP-01 challenge solver in cert-manager creates temporary Ingress resources to complete ACME challenges — these will use whichever IngressClass you configure in your Issuer or ClusterIssuer. If you migrate to Gateway API, cert-manager has added Gateway API support (HTTPRoute-based HTTP-01 challenges) starting from version 1.14. DNS-01 challenges are entirely unaffected by controller choice. Update your Issuer configuration to reference the new IngressClass during migration.

How severe is the performance difference between ingress-nginx and alternatives?

For the vast majority of production workloads, the performance difference between mature controllers (ingress-nginx, Traefik, HAProxy, Envoy) is not the deciding factor. All of them can handle tens of thousands of requests per second on reasonable hardware, and the bottleneck is typically the backend services, not the ingress layer. The notable exception is Cilium with eBPF-based forwarding, which eliminates userspace proxy overhead entirely and can show measurable latency reduction at high percentiles for latency-sensitive workloads. If you are running at a scale where ingress controller throughput is actually the constraint, you already have the engineering resources to benchmark your specific workload profile against candidate controllers before committing.

Should we just move everything to a cloud provider’s managed load balancer and skip the in-cluster controller?

This is a legitimate option for teams on managed Kubernetes (EKS, GKE, AKS). Cloud-native load balancers (AWS ALB via AWS Load Balancer Controller, GKE Gateway, Azure Application Gateway Ingress Controller) eliminate the operational burden of managing an in-cluster controller and integrate deeply with cloud IAM, WAF, and observability services. The trade-offs are cost (cloud LBs charge per rule and per hour), vendor lock-in, and reduced portability. For purely cloud-native workloads with no multi-cloud or on-premises requirements, cloud-managed load balancers are worth serious consideration and sidestep the ingress-nginx problem entirely. For hybrid or multi-cluster environments, in-cluster controllers maintain an advantage in consistency and portability.