Ingresses have been, since the early versions of Kubernetes, the most common way to expose applications to the outside. Although their initial design was simple and elegant, the success of Kubernetes and the growing complexity of use cases have turned Ingress into a problematic piece: limited, inconsistent between vendors, and difficult to govern in enterprise environments.
In this article, we analyze why Ingresses have become a constant source of friction, how different Ingress Controllers have influenced this situation, and why more and more organizations are considering alternatives like Gateway API.
What Ingresses are and why they were designed this way
The Ingress ecosystem revolves around two main resources:
🏷️ IngressClass
Defines which controller will manage the associated Ingresses. Its scope is cluster-wide, so it is usually managed by the platform team.
🌐 Ingress
It is the resource that developers use to expose a service. It allows defining routes, domains, TLS certificates, and little more.
Its specification is minimal by design, which allowed for rapid adoption, but also laid the foundation for current problems.
The problem: a standard too simple for complex needs
As Kubernetes became an enterprise standard, users wanted to replicate advanced configurations of traditional proxies: rewrites, timeouts, custom headers, CORS, etc. But Ingress did not provide native support for all this.
Vendors reacted… and chaos was born.
Annotations vs CRDs: two incompatible paths
Different Ingress Controllers have taken very different paths to add advanced capabilities:
📝 Annotations (NGINX, HAProxy…)
Advantages:
Flexible and easy to use
Directly in the Ingress resource
Disadvantages:
Hundreds of proprietary annotations
Fragmented documentation
Non-portable configurations between vendors
📦 Custom CRDs (Traefik, Kong…)
Advantages:
More structured and powerful
Better validation and control
Disadvantages:
Adds new non-standard objects
Requires installation and management
Less interoperability
Result? Infrastructures deeply coupled to a vendor, complicating migrations, audits, and automation.
The complexity for development teams
The design of Ingress implies two very different responsibilities:
Platform: defines IngressClass
Application: defines Ingress
But the reality is that the developer ends up making decisions that should be the responsibility of the platform area:
Certificates
Security policies
Rewrite rules
CORS
Timeouts
Corporate naming practices
This causes:
Inconsistent configurations
Bottlenecks in reviews
Constant dependency between teams
Lack of effective standardization
In large companies, where security and governance are critical, this is especially problematic.
NGINX Ingress: the decommissioning that reignited the debate
The recent decommissioning of the NGINX Ingress Controller has highlighted the fragility of the ecosystem:
This has reignited the conversation about the need for a real standard… and there appears Gateway API.
Gateway API: a promising alternative (but not perfect)
Gateway API was born to solve many of the limitations of Ingress:
Clear separation of responsibilities (infrastructure vs application)
Standardized extensibility
More types of routes (HTTPRoute, TCPRoute…)
Greater expressiveness without relying on proprietary annotations
But it also brings challenges:
Requires gradual adoption
Not all vendors implement the same
Migration is not trivial
Even so, it is shaping up to be the future of traffic management in Kubernetes.
Conclusion
Ingresses have been fundamental to the success of Kubernetes, but their own simplicity has led them to become a bottleneck. The lack of interoperability, differences between vendors, and complex governance in enterprise environments make it clear that it is time to adopt more mature models.
Gateway API is not perfect, but it moves in the right direction. Organizations that want future stability should start planning their transition.
What is Kubernetes Node Affinity? Benefits and Core Concepts
Kubernetes node affinity is an essential scheduling feature that allows you to control pod placement based on node labels and properties. By using node affinity rules, you can specify constraints on which nodes pods can be scheduled, enabling you to optimize resource allocation and enhance performance.
Node affinity works by allowing you to define rules for pod scheduling based on node labels. When defining node affinity rules, you have two options: required and preferred rules. Required rules ensure that pods are scheduled only on nodes that satisfy the defined criteria. If no suitable node is available, the pod remains unscheduled. On the other hand, preferred rules provide a soft constraint and attempt to schedule pods on nodes that match the specified criteria. However, if no such node is available, the pod can still be scheduled on other nodes.
Node affinity rules are an “expanded” option of the simply way by using node selectors. Node selectors are a simple form of node affinity that allows you to assign labels to nodes and match those labels with selectors defined in the pod specification. By specifying a node selector, you can ensure that pods are scheduled only on nodes with matching labels. Node selectors are useful for basic affinity requirements but lack the flexibility and fine-grained control provided by more advanced affinity options.
Node Affinity Trade-offs: Required vs Preferred Rules and Failure Scenarios
But this awesome capability has some trade-offs that you need to take in consideration because nothing comes with a price that you need to be aware of, so, let’s go to the important question, what is the worst case scenario of using any of those options?
Consider a stateful workload, like a distributed database (e.g., etcd or ZooKeeper), deployed with three replicas for consensus and fault tolerance. So you decide to define a set of nodes for this workload and use node affinity rules to ensure the pods are scheduled to those nodes. And, you need to think: should I use the preferred mode or the requiredMode?
Let’s say that you go with the required option and you define it like this, what happen if one of your nodes goes down? The pod will be try to be rescheduled again and unless there are another node “with same label” to that, it cannot be deployed? If you additional defined a pod anti-affinity rule to ensure each of the replicas is in a different host to ensure that in case that one node is going down you lose only a single replica, you’re losing the option to rescheudle the workload even if you have another nodes without the label available. So, you’re not in a so reliable option.
Ok, so you go with the preferred to ensure that you workload is for sure scheduled even if it is in another node, and in that case you can end up on the situation that those nodes are scheduled on other nodes keeping those nodes with the proper label without the workload that they should have, making the situation strange and more difficult to administer because you cannot ensure your workloads is on the nodes that you expected to be.
Additional to that, if the nodes has even taints to ensure other workloads cannot be placed there, you can end up in a situation that the “labeled-pods” are scheduled on non-labeled nodes, and the non-labeled pods cannot use the nodes because they’re tainted and can be not be able to use the un-labeled ones if there are not enough resources. So you’re generating an impact on the other workloasd and potentially affecting the schedulling of the other workloads.
Preparing for Unexpected Outages with Node Affinity
So, as you can see, each decision has some disadvatanges that you need to take in consdieration before defining those rules, because if you don’t, you will figure it out when this happen on an production enviornment probably as a result of some unexpected outage, because we all know that in the meantime that nothing bad happens everything works as expected, but the potential of these solutions and its reason to be used is exactly to provide the tools and the options to be prepared when bad things happens.
So, next time that you need to define a node affinity rule try to think about the disadvantages of each of the option and try to select that one that works best for you and mitigate the problems that it can bring to the table of your production environment.
What is the difference between nodeSelector and node affinity in Kubernetes?
nodeSelector is a simple field that requires a node to have all specified labels. Node affinity is a more expressive API that supports complex operators like In, NotIn, and Exists, and distinguishes between hard (requiredDuringScheduling...) and soft (preferredDuringScheduling...) constraints. Use nodeSelector for basic needs; use node affinity for advanced scheduling logic.
When should I use required vs preferred node affinity rules?
Use required rules for strict placement needs, like licensing constraints or specific hardware (e.g., GPU nodes). Use preferred rules for optimization, like trying to place pods on nodes in the same availability zone for lower latency. Be aware that required rules can prevent scheduling during node failures, while preferred rules may not guarantee optimal placement.
What are the risks of using required node affinity?
The primary risk is scheduling failure. If no node matches the required rules (e.g., due to a failure or label mismatch), the pod will remain Pending. This can lead to application downtime, especially if combined with Pod Anti-Affinity, which further restricts eligible nodes. Always ensure you have enough labeled nodes to handle failures.
How does node affinity interact with taints and tolerations?
They work sequentially. First, the scheduler filters nodes based on node affinity/selector rules. Then, from the filtered nodes, it checks taints and tolerations. A pod will only be scheduled on a node that satisfies both its affinity/selector requirements and for which the pod has a matching toleration for all the node’s taints.
What are best practices for defining node affinity labels?
Use clear, descriptive label keys (e.g., node.kubernetes.io/instance-type, topology.kubernetes.io/zone). Prefer built-in labels where possible. Document the purpose of custom labels. Combine node affinity with pod anti-affinity carefully to avoid over-constraining the scheduler. Test scenarios with node failures.
As Kubernetes clusters become an integral part of infrastructure, maintaining compliance with security and configuration policies is crucial. Kyverno, a policy engine designed for Kubernetes, can be integrated into your CI/CD pipelines to enforce configuration standards and automate policy checks. In this article, we’ll walk through integrating Kyverno CLI with GitHub Actions, providing a seamless workflow for validating Kubernetes manifests before they reach your cluster.
What is Kyverno CLI?
Kyverno is a Kubernetes-native policy management tool, enabling users to enforce best practices, security protocols, and compliance across clusters. Kyverno CLI is a command-line interface that lets you apply, test, and validate policies against YAML manifests locally or in CI/CD pipelines. By integrating Kyverno CLI with GitHub Actions, you can automate these policy checks, ensuring code quality and compliance before deploying resources to Kubernetes.
Benefits of Using Kyverno CLI in CI/CD Pipelines
Integrating Kyverno into your CI/CD workflow provides several advantages:
Automated Policy Validation: Detect policy violations early in the CI/CD pipeline, preventing misconfigured resources from deployment.
Enhanced Security Compliance: Kyverno enables checks for security best practices and compliance frameworks.
Faster Development: Early feedback on policy violations streamlines the process, allowing developers to fix issues promptly.
Setting Up Kyverno CLI in GitHub Actions
Step 1: Install Kyverno CLI
To use Kyverno in your pipeline, you need to install the Kyverno CLI in your GitHub Actions workflow. You can specify the Kyverno version required for your project or use the latest version.
Here’s a sample GitHub Actions YAML configuration to install Kyverno CLI:
name: CI Pipeline with Kyverno Policy Checks
on:
push:
branches:
- main
pull_request:
branches:
- main
jobs:
kyverno-policy-check:
runs-on: ubuntu-latest
steps:
- name: Checkout Code
uses: actions/checkout@v2
- name: Install Kyverno CLI
run: |
curl -LO https://github.com/kyverno/kyverno/releases/download/v<version>/kyverno-cli-linux.tar.gz
tar -xzf kyverno-cli-linux.tar.gz
sudo mv kyverno /usr/local/bin/
Replace <version> with the version of Kyverno CLI you wish to use. Alternatively, you can replace it with latest to always fetch the latest release.
Step 2: Define Policies for Validation
Create a directory in your repository to store Kyverno policies. These policies define the standards that your Kubernetes resources should comply with. For example, create a directory structure as follows:
Each policy is defined in YAML format and can be customized to meet specific requirements. Below are examples of policies that might be used:
Disallow latest Tag in Images: Prevents the use of the latest tag to ensure version consistency.
Enforce CPU/Memory Limits: Ensures resource limits are set for containers, which can prevent resource abuse.
Step 3: Add a GitHub Actions Step to Validate Manifests
In this step, you’ll use Kyverno CLI to validate Kubernetes manifests against the policies defined in the .github/policies directory. If a manifest fails validation, the pipeline will halt, preventing non-compliant resources from being deployed.
Here’s the YAML configuration to validate manifests:
Replace manifests/ with the path to your Kubernetes manifests in the repository. This command applies all policies in .github/policies against each YAML file in the manifests directory, stopping the pipeline if any non-compliant configurations are detected.
Step 4: Handle Validation Results
To make the output of Kyverno CLI more readable, you can use additional GitHub Actions steps to format and handle the results. For instance, you might set up a conditional step to notify the team if any manifest is non-compliant:
- name: Check for Policy Violations
if: failure()
run: echo "Policy violation detected. Please review the failed validation."
Alternatively, you could configure notifications to alert your team through Slack, email, or other integrations whenever a policy violation is identified.
—
Example: Validating a Kubernetes Manifest
Suppose you have a manifest defining a Kubernetes deployment as follows:
The policy disallow-latest-tag.yaml checks if any container image uses the latest tag and rejects it. When this manifest is processed, Kyverno CLI flags the image and halts the CI/CD pipeline with an error, preventing the deployment of this manifest until corrected.
Conclusion
Integrating Kyverno CLI into a GitHub Actions CI/CD pipeline offers a robust, automated solution for enforcing Kubernetes policies. With this setup, you can ensure Kubernetes resources are compliant with best practices and security standards before they reach production, enhancing the stability and security of your deployments.
Introduction OpenShift, Red Hat’s Kubernetes platform, has its own way of exposing services to external clients. In vanilla Kubernetes, you would typically use an Ingress resource along with an ingress controller to route external traffic to services. OpenShift, however, introduced the concept of a Route and an integrated Router (built on HAProxy) early on, before Kubernetes Ingress even existed. Today, OpenShift supports both Routes and standard Ingress objects, which can sometimes lead to confusion about when to use each and how they relate.
This article explores how OpenShift handles Kubernetes Ingress resources, how they translate to Routes, the limitations of this approach, and guidance on when to use Ingress versus Routes.
OpenShift Routes and the Router: A Quick Overview
OpenShift Routes are OpenShift-specific resources designed to expose services externally. They are served by the OpenShift Router, which is an HAProxy-based proxy running inside the cluster. Routes support advanced features such as:
Because Routes are OpenShift-native, the Router understands these features natively and can be configured accordingly. This tight integration enables powerful and flexible routing capabilities tailored to OpenShift environments.
Using Kubernetes Ingress in OpenShift (Default Behavior)
Starting with OpenShift Container Platform (OCP) 3.10, Kubernetes Ingress resources are supported. When you create an Ingress, OpenShift automatically translates it into an equivalent Route behind the scenes. This means you can use standard Kubernetes Ingress manifests, and OpenShift will handle exposing your services externally by creating Routes accordingly.
This automatic translation simplifies migration and supports basic use cases without requiring Route-specific manifests.
Tuning Behavior with Annotations (Ingress ➝ Route)
When you use Ingress on OpenShift, only OpenShift-aware annotations are honored during the Ingress ➝ Route translation. Controller-specific annotations for other ingress controllers (e.g., nginx.ingress.kubernetes.io/*) are ignored by the OpenShift Router. The following annotations are commonly used and supported by the OpenShift router to tweak the generated Route:
Purpose
Annotation
Typical Values
Effect on Generated Route
TLS termination
route.openshift.io/termination
edge · reencrypt · passthrough
Sets Route spec.tls.termination to the chosen mode.
This Ingress will be realized as a Route with edge TLS and an automatic HTTP→HTTPS redirect, using least connections balancing and a 60s route timeout. The HSTS header will be added by the router on HTTPS responses.
Limitations of Using Ingress to Generate Routes While convenient, using Ingress to generate Routes has limitations:
Missing advanced features: Weighted backends and sticky sessions require Route-specific annotations and are not supported via Ingress.
TLS passthrough and re-encrypt modes: These require OpenShift-specific annotations on Routes and are not supported through standard Ingress.
Ingress without host: An Ingress without a hostname will not create a Route; Routes require a host.
Wildcard hosts: Wildcard hosts (e.g., *.example.com) are only supported via Routes, not Ingress.
Annotation compatibility: Some OpenShift Route annotations do not have equivalents in Ingress, leading to configuration gaps.
Protocol support: Ingress supports only HTTP/HTTPS protocols, while Routes can handle non-HTTP protocols with passthrough TLS.
Config drift risk: Because Routes created from Ingress are managed by OpenShift, manual edits to the generated Route may be overwritten or cause inconsistencies.
These limitations mean that for advanced routing configurations or OpenShift-specific features, using Routes directly is preferable.
When to Use Ingress vs. When to Use Routes Choosing between Ingress and Routes depends on your requirements:
Use Ingress if:
You want portability across Kubernetes platforms.
You have existing Ingress manifests and want to minimize changes.
Your application uses only basic HTTP or HTTPS routing.
You prefer platform-neutral manifests for CI/CD pipelines.
Use Routes if:
You need advanced routing features like weighted backends, sticky sessions, or multiple TLS termination modes.
Your deployment is OpenShift-specific and can leverage OpenShift-native features.
You require stability and full support for OpenShift routing capabilities.
You need to expose non-HTTP protocols or use TLS passthrough/re-encrypt modes.
You want to use wildcard hosts or custom annotations not supported by Ingress.
In many cases, teams use a combination: Ingress for portability and Routes for advanced or OpenShift-specific needs.
Conclusion
On OpenShift, Kubernetes Ingress resources are automatically converted into Routes, enabling basic external service exposure with minimal effort. This allows users to leverage existing Kubernetes manifests and maintain portability. However, for advanced routing scenarios and to fully utilize OpenShift’s powerful Router features, using Routes directly is recommended.
Both Ingress and Routes coexist seamlessly on OpenShift, allowing you to choose the right tool for your application’s requirements.
What is the difference between an OpenShift Route and a Kubernetes Ingress?
Route is OpenShift’s native, pre-Ingress API for exposing services; Ingress is the upstream Kubernetes standard. Functionally they overlap, but Routes support OpenShift-specific TLS modes (edge, reencrypt, passthrough) natively, while Ingress is portable across any Kubernetes distribution.
Does OpenShift support standard Kubernetes Ingress?
Yes. When you create an Ingress, the OpenShift router generates the equivalent Route objects automatically, so upstream manifests work unmodified. The generated Routes are owned by the Ingress — edit the Ingress, not the Routes.
Should I use Route or Ingress in OpenShift?
Use Ingress if your manifests must stay portable across distributions or come from upstream Helm charts. Use Route when you need OpenShift-specific behavior (reencrypt TLS, passthrough) or your platform is OpenShift-only by policy. Mixing both works, but standardising on one keeps GitOps diffs sane.
How does TLS work when an Ingress generates a Route?
TLS config in the Ingress (a tls block with a certificate Secret) maps to an edge-terminated Route. For reencrypt or passthrough you need either Route objects directly or the OpenShift-specific annotation route.openshift.io/termination on the Ingress.
Every Kubernetes cluster runs on Linux. But the distribution you choose for your nodes determines how much time you spend patching, hardening, debugging SSH sessions, and dealing with configuration drift across your fleet. General-purpose distributions like Ubuntu and Debian were designed to run anything: web servers, desktops, databases, and yes, Kubernetes. That flexibility is also their biggest liability when your only job is running containers.
Talos Linux takes a radically different approach. It strips away everything a Kubernetes node does not need: there is no shell, no SSH daemon, no package manager, and no way to log in interactively. The entire operating system is managed through an API, and every change is declarative. If that sounds extreme, it is. But it solves real problems that traditional distributions cannot address without layers of additional tooling.
This guide is a comprehensive deep dive into Talos Linux: what it is, how its architecture works, how it compares to alternatives like Flatcar and Bottlerocket, how to install and operate it, and when you should (and should not) use it. Whether you are evaluating Talos for a production fleet or a homelab, this is everything you need to make an informed decision.
What Is Talos Linux
Talos Linux is a minimal, immutable operating system designed exclusively to run Kubernetes. It is developed by Sidero Labs and distributed as a single system image that boots into a Kubernetes-ready state. There is no general-purpose userland. No bash shell. No ability to SSH into a node and run commands. Every aspect of machine configuration — from network settings to Kubernetes component flags — is expressed in a YAML document called the machine config and applied through an authenticated gRPC API.
The core design principles are:
Immutable — The root filesystem is read-only and mounted from a SquashFS image. You cannot install packages, modify system binaries, or alter the OS at runtime.
API-driven — All management happens through talosctl, a CLI that communicates with the Talos API over mutual TLS. There is no SSH and no interactive console.
Minimal — The OS ships only what Kubernetes needs: a Linux kernel, containerd, the kubelet, etcd (on control plane nodes), and the Talos machinery. The installed image is roughly 80 MB.
Declarative — The desired machine state is defined in a YAML config. Applying a new config converges the node to the desired state, similar to how Kubernetes reconciles workloads.
Secure by default — No shell access means no attack vector through compromised credentials. All API communication requires mutual TLS authentication. The attack surface is drastically smaller than any traditional distribution.
Talos supports bare metal, VMware vSphere, AWS, Azure, GCP, Hetzner, Equinix Metal, Oracle Cloud, and several other platforms. It also runs on single-board computers like Raspberry Pi and NVIDIA Jetson, making it viable for edge deployments. For a broader perspective on how immutable infrastructure fits into the Kubernetes ecosystem, see our Kubernetes security best practices guide.
Architecture Deep Dive
Understanding Talos at an architectural level is essential before deploying it. The design choices are unconventional compared to what most Linux administrators expect, and they explain both its strengths and its constraints.
The machined Daemon and API-Driven Management
At the heart of Talos is machined, a single PID-1 process that replaces systemd, init, and every other service manager. When a Talos node boots, machined starts, reads its machine configuration, and orchestrates the entire lifecycle: networking, disk setup, containerd, the kubelet, and etcd (on control plane nodes).
machined exposes a gRPC API over port 50000 (for the trustd/machine API) and port 50001 (for the maintenance API during initial provisioning). This is the only way to interact with the node. The talosctl CLI is the primary client, authenticating with mutual TLS certificates generated during cluster bootstrapping.
Key API operations include:
talosctl apply-config — Push a new or updated machine configuration.
talosctl upgrade — Trigger an in-place OS upgrade.
talosctl dmesg — Stream kernel messages in real time.
talosctl logs — Read logs from any Talos service (etcd, kubelet, containerd).
talosctl get — Inspect resource state (network interfaces, disks, services).
talosctl reset — Wipe a node and return it to maintenance mode.
This API-first model eliminates configuration drift by design. There is no way for an operator to SSH into a node, run an ad-hoc command, and leave the system in an undocumented state. Every change flows through the same declarative path.
System Partitions Layout
Talos partitions the disk into a well-defined layout that separates immutable system data from mutable state:
Partition
Purpose
Mutable
EFI
EFI System Partition for UEFI boot
No
BIOS
BIOS boot partition (legacy boot)
No
BOOT
Contains the kernel and initramfs
No (replaced during upgrades)
META
Stores metadata like machine UUID and upgrade status
Limited
STATE
Holds the machine configuration and PKI material
Yes (managed by machined)
EPHEMERAL
Mounted at /var, stores containerd images, kubelet data, etcd data, and pod logs
Yes (wiped on reset)
The STATE partition is critical: it persists the machine config and TLS certificates across reboots and upgrades. The EPHEMERAL partition holds everything that can be reconstructed — container images, pod volumes (emptyDir), and etcd data on control plane nodes. When you run talosctl reset, the EPHEMERAL partition is wiped, but STATE can optionally be preserved.
This layout means that an OS upgrade replaces the BOOT partition contents (kernel + initramfs) while leaving your machine configuration and Kubernetes state untouched. If an upgrade fails, Talos rolls back to the previous BOOT image automatically.
Boot Process and Kubernetes Bootstrapping
The Talos boot sequence is deterministic and fast, typically completing in under 60 seconds on modern hardware:
Firmware → Bootloader — UEFI or BIOS loads GRUB, which loads the Talos kernel and initramfs.
Kernel init → machined — The kernel starts machined as PID 1. There is no init system in between.
Machine config discovery — machined checks the STATE partition for an existing config. If none is found (first boot), it enters maintenance mode and listens on the maintenance API for a config to be applied.
Network configuration — Networking is brought up based on the machine config (DHCP or static).
Disk setup — Partitions are created or validated. The EPHEMERAL partition is formatted if missing.
containerd starts — The container runtime is launched.
etcd starts (control plane only) — etcd is started and joins the existing cluster, or waits for a bootstrap command.
kubelet starts — The kubelet registers the node with the Kubernetes API server.
The first control plane node requires a one-time bootstrap command (talosctl bootstrap) to initialize the etcd cluster and generate the Kubernetes control plane static pods. Subsequent control plane nodes join automatically.
Security Model: No SSH, Mutual TLS, API-Only
Talos Linux implements a zero-trust security model at the OS level. Every API request is authenticated using mutual TLS (mTLS). When you generate a cluster configuration with talosctl gen config, it produces a Certificate Authority (CA) that signs both the client (operator) and server (node) certificates.
The security implications are significant:
No shell access — There is no /bin/sh, no /bin/bash, no login capability. Even if an attacker gains network access to the node, there is no shell to exploit.
No SSH daemon — Port 22 is not open. There is no sshd binary on the system.
No package manager — You cannot install tools, backdoors, or persistence mechanisms on the host.
Read-only rootfs — Even with theoretical root access, the filesystem cannot be modified.
Mutual TLS everywhere — The Talos API, etcd communication, and inter-node trust all use mTLS. Certificates can be rotated without downtime.
This does not make Talos invulnerable — kernel exploits and container escape vulnerabilities still apply. But it eliminates the most common attack vectors in Kubernetes node compromise: SSH credential theft, unauthorized package installation, and persistent rootkits.
Talos Linux vs Alternatives: Comparison Table
Choosing a node OS depends on your operational model, cloud provider, and team experience. Here is how Talos Linux compares to the most common alternatives for Kubernetes node operating systems.
Feature
Talos Linux
Ubuntu / Debian
Flatcar Container Linux
Bottlerocket (AWS)
RancherOS / k3OS
Mutability
Fully immutable rootfs
Fully mutable
Immutable rootfs, writable /etc
Immutable rootfs
Mostly immutable
SSH Access
None (no sshd)
Yes (default)
Yes (default)
Optional (admin container)
Yes
Shell Access
None
Full shell
Full shell
Limited (via admin container)
Full shell
Management Model
Declarative API (gRPC)
Imperative (apt, SSH)
Declarative (Ignition) + SSH
Declarative (TOML settings API)
cloud-init + SSH
Update Mechanism
A/B image swap with rollback
apt upgrade (in-place)
A/B image swap (Nebraska/FLUO)
A/B image swap
Image swap
Container Runtime
containerd
containerd or CRI-O
containerd (Docker optional)
containerd
Docker (RancherOS), containerd (k3OS)
Kubernetes Integration
Built-in (kubelet, etcd bundled)
Manual (kubeadm, etc.)
Manual (kubeadm, etc.)
EKS-optimized
Built-in (k3s bundled)
Cloud Support
AWS, Azure, GCP, Hetzner, bare metal, VMware, and more
All clouds
AWS, Azure, GCP, bare metal, VMware
AWS only
Limited
Image Size
~80 MB
~1-2 GB
~300 MB
~200 MB
~150 MB
Config Drift
Impossible (API-only)
Common without tooling
Possible (SSH access)
Low (API + limited shell)
Possible
Talos Linux vs Ubuntu / Debian
Ubuntu and Debian are the default choices for most Kubernetes deployments, especially when using kubeadm or managed installers. They work. But they carry everything a general-purpose OS includes: a package manager, a full shell, hundreds of system services, and thousands of binaries that your Kubernetes nodes never use.
The operational burden is real: you need to patch the OS independently from Kubernetes, harden SSH, configure unattended upgrades, manage user accounts, and run CIS benchmarks to verify compliance. With Talos, these concerns disappear because the attack surface simply does not exist. The trade-off is that you lose the ability to SSH in and debug problems the traditional way.
Talos Linux vs Flatcar Container Linux
Flatcar Container Linux (the successor to CoreOS Container Linux) is the closest philosophical match to Talos. Both use immutable root filesystems and image-based updates. However, Flatcar retains SSH access and a full shell, which means an operator can still log in and make ad-hoc changes. Flatcar uses Ignition for initial provisioning and systemd for service management.
The key difference is that Flatcar is a container-optimized general-purpose OS, while Talos is a Kubernetes-only OS. Flatcar can run arbitrary containers and system services. Talos runs only Kubernetes. If you need SSH as a safety net during your transition to immutable infrastructure, Flatcar is a pragmatic middle ground. If you want to enforce immutability with no escape hatches, Talos is the stronger choice.
Talos Linux vs Bottlerocket
Bottlerocket is AWS’s purpose-built container OS, designed for EKS and ECS. Like Talos, it has an immutable rootfs and an API-driven settings model. Unlike Talos, it provides an optional “admin container” that gives you a shell for debugging, and it is heavily optimized for the AWS ecosystem.
If you run exclusively on AWS with EKS, Bottlerocket is the path of least resistance. If you need a multi-cloud or bare-metal solution with integrated Kubernetes bootstrapping, Talos is significantly more flexible. Bottlerocket also does not bootstrap Kubernetes itself — it relies on EKS or an external installer.
Talos Linux vs RancherOS / k3OS
RancherOS and k3OS were early attempts at minimal container-focused Linux distributions. RancherOS ran the entire system as Docker containers. k3OS bundled k3s (lightweight Kubernetes) into the OS. Both projects have been deprecated or are in maintenance mode. Talos is the actively developed, production-grade successor to this category. If you are currently running k3OS, Talos is the natural migration path.
Installation and Cluster Bootstrap
Setting up a Talos cluster follows a consistent workflow regardless of the platform: generate configs, boot nodes, apply configs, bootstrap. Here is a step-by-step walkthrough.
Step 1: Install talosctl
Download the talosctl binary for your platform. On macOS with Homebrew:
brew install siderolabs/tap/talosctl
On Linux:
curl -sL https://talos.dev/install | sh
Step 2: Generate Machine Configurations
The talosctl gen config command generates a full set of machine configurations: one for control plane nodes, one for workers, and a talosconfig file containing the client credentials.
talosctl gen config my-cluster https://10.0.0.10:6443
--output-dir _out
This creates three files in the _out directory:
controlplane.yaml — Machine config for control plane nodes.
worker.yaml — Machine config for worker nodes.
talosconfig — Client configuration with the CA certificate and client key for mTLS authentication.
The endpoint URL (https://10.0.0.10:6443) should point to the Kubernetes API server address — either a load balancer VIP or the IP of your first control plane node.
Step 3: Boot Nodes with Talos
How you boot depends on the platform:
Bare metal — Write the Talos ISO or disk image to a USB drive or PXE boot. The node boots into maintenance mode, waiting for a config.
VMware — Deploy the OVA template, or use the ISO in a VM. Talos provides official OVA images.
AWS — Use the official Talos AMI. Launch EC2 instances with the AMI and pass the machine config as user-data.
Azure / GCP — Use the official images from Sidero Labs’ image factory. Pass the machine config through the platform’s metadata service.
Step 4: Apply Configuration and Bootstrap
Once nodes are booted and in maintenance mode, apply the machine configs:
# Configure talosctl to use the generated credentials
export TALOSCONFIG="_out/talosconfig"
# Apply config to the first control plane node
talosctl apply-config --insecure
--nodes 10.0.0.10
--file _out/controlplane.yaml
# Apply config to worker nodes
talosctl apply-config --insecure
--nodes 10.0.0.20
--file _out/worker.yaml
The --insecure flag is required for the initial config application because the node does not yet have TLS certificates. After the config is applied, all subsequent communication uses mTLS.
Now bootstrap the Kubernetes cluster from the first control plane node:
# Set the endpoint and node
talosctl config endpoint 10.0.0.10
talosctl config node 10.0.0.10
# Bootstrap etcd and the control plane
talosctl bootstrap
This command initializes etcd, generates the Kubernetes PKI, and starts the control plane static pods. Within a minute or two, the Kubernetes API server is available.
Step 5: Retrieve kubeconfig and Verify
# Get the kubeconfig
talosctl kubeconfig -n 10.0.0.10
# Verify the cluster
kubectl get nodes
kubectl get pods -A
Essential talosctl Commands
Once the cluster is running, these are the commands you will use daily:
# Check node health
talosctl health --nodes 10.0.0.10
# Stream kernel messages (equivalent to dmesg -w)
talosctl dmesg --nodes 10.0.0.10 --follow
# View service logs
talosctl logs kubelet --nodes 10.0.0.10
talosctl logs etcd --nodes 10.0.0.10
# List running services
talosctl services --nodes 10.0.0.10
# Get machine config (current running config)
talosctl get machineconfig --nodes 10.0.0.10
# Inspect resource state
talosctl get members --nodes 10.0.0.10
talosctl get addresses --nodes 10.0.0.10
Day-2 Operations
Installation is only the beginning. The real value of Talos emerges in day-2 operations: upgrades, config changes, and cluster maintenance. This is where the declarative, API-driven model pays dividends.
Upgrading Talos Linux
Talos upgrades are performed node by node through the API. The process downloads the new OS image, writes it to the inactive boot partition, and reboots the node into the new version. If the upgrade fails, the node automatically rolls back to the previous image.
# Upgrade a single node
talosctl upgrade --nodes 10.0.0.10
--image ghcr.io/siderolabs/installer:v1.9.0
# Upgrade with --preserve to keep the EPHEMERAL partition
talosctl upgrade --nodes 10.0.0.10
--image ghcr.io/siderolabs/installer:v1.9.0
--preserve
For production clusters, follow this sequence: upgrade control plane nodes one at a time, verify etcd health after each, then upgrade workers in a rolling fashion. The --preserve flag is important if you want to keep downloaded container images and avoid re-pulling everything after the reboot.
Upgrading Kubernetes Version
Kubernetes version upgrades are separate from Talos OS upgrades. You can run a newer version of Kubernetes on an older Talos release (within compatibility bounds). The upgrade is triggered through talosctl:
This command orchestrates the upgrade of all control plane components (kube-apiserver, kube-controller-manager, kube-scheduler, kube-proxy) and then rolls the kubelet version across all nodes. It respects PodDisruptionBudgets and cordons/drains nodes before upgrading.
Customizing Machine Config with Patches
As your cluster evolves, you will need to modify machine configurations — adding a registry mirror, changing kubelet flags, or configuring network bonding. Talos supports config patches that overlay changes onto the base config without replacing the entire file.
Patches can also be applied at generation time with talosctl gen config --config-patch, which is ideal for encoding environment-specific overrides into your GitOps pipeline.
etcd Management
Talos manages etcd as a first-class service, not as a manually deployed component. Common etcd operations are available through talosctl:
# Check etcd member list
talosctl etcd members --nodes 10.0.0.10
# Take an etcd snapshot (backup)
talosctl etcd snapshot db.snapshot --nodes 10.0.0.10
# Remove a failed etcd member
talosctl etcd remove-member --nodes 10.0.0.10
# Force a new etcd cluster from a single node (disaster recovery)
talosctl etcd forfeit-leadership --nodes 10.0.0.10
Regular etcd snapshots are non-negotiable for any production cluster. Automate this with a CronJob that calls the Talos API or runs talosctl etcd snapshot from an external host.
Limitations and When NOT to Use Talos Linux
Talos is not the right choice for every environment. Understanding its limitations is just as important as understanding its strengths.
No SSH Debugging
The most immediate pain point: when something goes wrong, you cannot SSH into the node and poke around. You are limited to what the Talos API exposes — logs, dmesg, service status, and resource state. For most Kubernetes issues, this is sufficient. But for low-level kernel or hardware debugging, you may need to boot the node from a different OS temporarily.
Talos does offer a talosctl dashboard command that provides a real-time TUI (text UI) showing CPU, memory, network, and service status. Combined with talosctl logs and talosctl dmesg, you can troubleshoot most problems. But the learning curve is real, especially for teams accustomed to reaching for htop and journalctl.
Learning Curve for Traditional Sysadmins
If your team manages infrastructure through SSH, Ansible playbooks, and shell scripts, Talos requires a fundamental shift in operational practices. There is no way to "just install" a debugging tool on a node. Everything must be done through the API or through Kubernetes workloads (DaemonSets with host-level access). This shift is valuable in the long run, but it requires investment in training and new workflows.
Custom Kernel Modules
Talos ships a specific kernel build with a curated set of modules. If your workload requires a custom kernel module (GPU drivers, specific storage drivers, or out-of-tree network drivers), you need to build a custom Talos image using the Talos image factory or the imager tool. This is supported but adds operational complexity compared to distributions where you can simply apt install a kernel module package.
Sidero Labs provides an Image Factory service that lets you build custom Talos images with additional system extensions (like NVIDIA drivers, iSCSI tools, or ZFS support) through a web interface or API.
Workloads Requiring Host-Level Access
Some workloads expect to interact with the host OS directly: log collectors that read /var/log, monitoring agents that read /proc, or security tools that install kernel modules. Most of these work in Talos (containerd's runtime allows host path mounts), but some assume a traditional Linux userland that simply does not exist. Evaluate your specific stack before committing.
Real-World Use Cases
Homelab and Learning
Talos is an excellent choice for homelab Kubernetes clusters. It runs on Raspberry Pi 4/5, Intel NUCs, and old laptops. The entire OS fits in minimal storage, and the declarative config model means you can rebuild your cluster from scratch in minutes by reapplying your machine configs. Many homelab operators use Talos with ArgoCD or Flux for a fully GitOps-managed stack.
Edge and Retail
Edge deployments benefit from Talos's small footprint, immutable design, and remote management. A retail chain with 500 store locations running local Kubernetes clusters can manage every node through the Talos API without ever needing physical or SSH access. The A/B upgrade mechanism ensures that a bad update does not brick a remote device.
Production Multi-Cloud Clusters
Talos provides a consistent node OS across AWS, Azure, GCP, and bare metal. This is valuable for organizations that run Kubernetes on multiple providers and want a single operational model for node management. Instead of maintaining separate AMIs, Azure images, and GCP images with different toolchains, you maintain one set of Talos machine configs with platform-specific patches.
Security-Sensitive Environments
For regulated industries (finance, healthcare, government), Talos's security posture simplifies compliance. The absence of SSH, shell, and package management eliminates entire categories of CIS benchmark requirements. Audit teams appreciate that there is no way for a rogue operator to install unauthorized software on the node OS. The immutable image model also simplifies forensics: if the OS hash does not match the known-good image, the node has been tampered with.
Frequently Asked Questions
Can you SSH into Talos Linux?
No. Talos Linux does not include an SSH daemon, a shell, or any interactive login mechanism. All node management is performed through the Talos API using talosctl. This is a deliberate design decision to eliminate the attack surface associated with shell access and prevent configuration drift from ad-hoc changes.
Is Talos Linux free and open source?
Yes. Talos Linux is open source under the Mozilla Public License 2.0. It is developed by Sidero Labs, which also offers Omni — a commercial SaaS platform for managing Talos clusters at scale. The OS itself is fully free to use in production without restrictions.
How do you debug a Talos Linux node without shell access?
Talos provides several debugging tools through its API: talosctl dmesg for kernel messages, talosctl logs <service> for service logs, talosctl dashboard for a real-time system overview, and talosctl get for inspecting resource state (network, disks, services). For deeper debugging, you can run a privileged DaemonSet pod with nsenter to access the host namespace from within Kubernetes.
Can Talos Linux run workloads other than Kubernetes?
No. Talos Linux is purpose-built exclusively for Kubernetes. It does not support running arbitrary containers, system services, or applications outside of the Kubernetes workload model. If you need to run non-Kubernetes workloads on the same host, consider Flatcar Container Linux or a traditional distribution.
What happens if a Talos upgrade fails?
Talos uses an A/B partition scheme for upgrades. The new image is written to the inactive boot partition, and the node reboots into it. If the new image fails to boot successfully (the health check does not pass within the configured timeout), the bootloader automatically reverts to the previous working image on the next reboot. This makes upgrades inherently safe and reversible without manual intervention.
Helm has long been the standard for managing Kubernetes applications using packaged charts, bringing a level of reproducibility and automation to the deployment process. However, some operational tasks, such as renaming a release or migrating objects between charts, have traditionally required cumbersome workarounds. With the introduction of the --take-ownership flag in Helm v3.17 (released in January 2025), a long-standing pain point is finally addressed—at least partially.
The take-ownership feature represents the continuing evolution of Helm. Learn about this and other cutting-edge capabilities in our Helm Charts Package Management Guide
In this post, we will explore:
What the --take-ownership flag does
Why it was needed
The caveats and limitations
Real-world use cases where it helps
When not to use it
Understanding Helm Release Ownership and Object Management
When Helm installs or upgrades a chart, it injects metadata—labels and annotations—into every managed Kubernetes object. These include:
This metadata serves an important role: Helm uses it to track and manage resources associated with each release. As a safeguard, Helm does not allow another release to modify objects it does not own and when you trying that you will see messages like the one below:
Error: Unable to continue with install: Service "provisioner-agent" in namespace "test-my-ns" exists and cannot be imported into the current release: invalid ownership metadata; annotation validation error: key "meta.helm.sh/release-name" must equal "dp-core-infrastructure11": current value is "dp-core-infrastructure"
While this protects users from accidental overwrites, it creates limitations for advanced use cases.
Why --take-ownership Was Needed
Let’s say you want to:
Rename an existing Helm release from api-v1 to api.
Move a ConfigMap or Service from one chart to another.
Rebuild state during GitOps reconciliation when previous Helm metadata has drifted.
Previously, your only option was to:
Uninstall the existing release.
Reinstall under the new name.
This approach introduces downtime, and in production systems, that’s often not acceptable.
You’re refactoring a large chart into smaller, modular ones and need to reassign certain Service or Secret objects.
This flag allows the new release to take control of the object without deleting or recreating it.
✅ 3. GitOps Drift Reconciliation
If objects were deployed out-of-band or their metadata changed unintentionally, GitOps tooling using Helm can recover without manual intervention using --take-ownership.
Best Practices and Recommendations
Use this flag intentionally, and document where it’s applied.
If possible, remove the previous release after migration to avoid confusion.
Monitor Helm’s behavior closely when managing shared objects.
For non-Helm-managed resources, continue to use kubectl annotate or kubectl label to manually align metadata.
Conclusion
The --take-ownership flag is a welcomed addition to Helm’s CLI arsenal. While not a universal solution, it smooths over many of the rough edges developers and SREs face during release evolution and GitOps adoption.
It brings a subtle but powerful improvement—especially in complex environments where resource ownership isn’t static.
Stay updated with Helm releases, and consider this flag your new ally in advanced release engineering.
The Modern Way: helm upgrade –take-ownership
Since Helm 3.17, most of the manual adoption workflow below is one flag:
With --take-ownership, Helm skips the ownership check that normally fails with “invalid ownership metadata; annotation validation error” and rewrites the ownership metadata to point at the new release. The resources’ spec is untouched — only the Helm bookkeeping changes. Two caveats: it applies to every conflicting resource in the chart (there is no per-resource opt-in), and it will happily steal resources from another Helm release, not just from kubectl-created objects — so run it with --dry-run first in anything shared.
Manual Adoption: The Three Ownership Markers
On older Helm versions (or when you want per-resource control), adoption means setting the three markers Helm checks before managing a resource:
After those three, the next helm upgrade treats the resource as its own and reconciles it against the chart template. This is the route to migrate kubectl-managed or kustomize-managed infrastructure into a chart incrementally, one resource at a time, without deleting anything.
Frequently Asked Questions
What does the Helm u002du002dtake-ownership flag do?
The u003ccodeu003eu002du002dtake-ownershipu003c/codeu003e flag allows Helm to bypass ownership validation and claim control of Kubernetes resources that belong to another release. It updates the u003cstrongu003emeta.helm.sh/release-nameu003c/strongu003e annotation to associate objects with the current release, enabling zero-downtime release renames and chart migrations.
When should I use Helm take ownership?
Use u003ccodeu003eu002du002dtake-ownershipu003c/codeu003e when renaming releases without downtime, migrating objects between charts, or fixing GitOps drift. It’s ideal for u003cstrongu003eproduction environmentsu003c/strongu003e where uninstall/reinstall cycles aren’t acceptable. Always document usage and clean up previous releases afterward.
What are the limitations of Helm take ownership?
The flag u003cstrongu003edoesn’t clean upu003c/strongu003e references from previous releases or protect against future uninstalls of the original release. It only works with Helm-managed resources, not completely unmanaged Kubernetes objects. Manual cleanup of old releases is still required.
Is Helm take ownership safe for production use?
Yes, but use it u003cstrongu003eintentionally and carefullyu003c/strongu003e. The flag bypasses Helm’s safety checks, so ensure you understand the ownership implications. Test in staging first, document all usage, and monitor for conflicts. Remove old releases after successful migration to avoid confusion.
Which Helm version introduced the take ownership flag?
The u003ccodeu003eu002du002dtake-ownershipu003c/codeu003e flag was introduced in u003cstrongu003eHelm v3.17u003c/strongu003e, released in January 2025. This feature addresses long-standing pain points with release renaming and chart migrations that previously required downtime-inducing uninstall/reinstall cycles.
In the Kubernetes ecosystem, security and governance are key aspects that need continuous attention. While Kubernetes offers some out-of-the-box (OOTB) security features such as Pod Security Admission (PSA), these might not be sufficient for complex environments with varying compliance requirements. This is where Kyverno comes into play, providing a powerful yet flexible solution for managing and enforcing policies across your cluster.
In this post, we will explore the key differences between Kyverno and PSA, explain how Kyverno can be used in different use cases, and show you how to install and deploy policies with it. Although custom policy creation will be covered in a separate post, we will reference some pre-built policies you can use right away.
What Is Kyverno?
Kyverno is a policy engine for Kubernetes that lets you validate, mutate and generate cluster resources using policies written as plain YAML — no new programming language required. It runs as an admission controller: when someone applies a resource, Kyverno intercepts the request and either allows it, rejects it, silently fixes it, or creates additional resources alongside it, according to the rules you have defined.
The name means “govern” in Greek, and that is a fair description of the job. Typical uses are enforcing that every pod sets resource limits, blocking images from untrusted registries, requiring specific labels on namespaces, automatically injecting a NetworkPolicy into every new namespace, or verifying image signatures before a workload is admitted.
Three properties are worth knowing before you evaluate it:
Policies are Kubernetes resources. A Kyverno policy is a ClusterPolicy or Policy object written in YAML, so it goes through the same review, GitOps and RBAC machinery as the rest of your manifests.
It does more than say no. Validation is only one of three modes. Kyverno can also mutate (patch a resource on the way in) and generate (create related resources automatically), which is where most of its day-to-day value ends up coming from.
It is Kubernetes-only by design. That is a deliberate trade-off, and it is the main axis on which it differs from OPA Gatekeeper — see the comparison below.
The current release at the time of writing is Kyverno v1.18.2 (July 2026), and recent versions support both YAML-based and CEL-based policy expressions.
Kyverno vs OPA Gatekeeper: Which Policy Engine?
This is the comparison most teams actually need to make, and after several years of both projects converging it is rarely about raw capability. The short version: choose Kyverno if your policies are Kubernetes-only and you want them in YAML; choose OPA Gatekeeper if you need maximum expressiveness or want to reuse the same policy language outside Kubernetes.
Kyverno
OPA Gatekeeper
Policy language
YAML (plus CEL); nothing new to learn
Rego, via ConstraintTemplates and Constraints
Learning curve
Low — it looks like the manifests you already write
Higher — Rego is a genuine language to learn
Validate
✅ first-class
✅ first-class, its strongest area
Mutate
✅ first-class policy type
✅ added later, via dedicated resources rather than Rego
Generate resources
✅ native
❌ not a design goal
Scope
Kubernetes only
Kubernetes, plus APIs, microservices, Terraform — same Rego everywhere
Complex programmatic logic
Workable, but YAML shows its limits
Where Rego pulls ahead
CNCF maturity
Graduated
Graduated (part of Open Policy Agent)
Current release
v1.18.2 (Jul 2026)
v3.23.0 (Jul 2026)
There is now a third option that did not exist when this debate started: ValidatingAdmissionPolicy with CEL, built into Kubernetes itself. It needs no extra components at all, so if a rule can be expressed in CEL it is worth reaching for first. The limitation is that CEL cannot make external calls or hold state — so image-signature verification, cross-resource lookups, mutation and generation still belong to Kyverno or Gatekeeper. In practice many clusters end up with native CEL policies for the simple rules and Kyverno for everything else.
Who Maintains Kyverno? Governance and Commercial Support
Kyverno is a CNCF Graduated project — the same maturity tier as Kubernetes, Prometheus and Envoy, and the highest the foundation awards. That matters for the question every platform team eventually asks: the project is not owned by a single vendor who can change the licence, and graduation requires a demonstrated governance model, security audit and a healthy contributor base.
It was created by Nirmata, who remain the most active contributor and offer commercial support and long-term support around it. If you need a support contract, an LTS commitment or consulting, Nirmata is the primary vendor, and several others also provide commercial Kyverno support — Giant Swarm, InfraCloud, BlakYaks and Kodekloud among them. The open-source project itself is free and Apache-2.0, and nothing in the upstream distribution is gated behind a commercial tier.
What is Pod Security Admission (PSA)?
Kubernetes introduced Pod Security Admission (PSA) as a replacement for the now deprecated PodSecurityPolicy (PSP). PSA focuses on enforcing three predefined levels of security: Privileged, Baseline, and Restricted. These levels control what pods are allowed to run in a namespace based on their security context configurations.
Privileged: Minimal restrictions, allowing privileged containers and host access.
Baseline: Applies standard restrictions, disallowing privileged containers and limiting host access.
Restricted: The strictest level, ensuring secure defaults and enforcing best practices for running containers.
While PSA is effective for basic security requirements, it lacks flexibility when enforcing fine-grained or custom policies. We have a full article covering this topic that you can read here.
Kyverno vs. PSA: Key Differences
Kyverno extends beyond the capabilities of PSA by offering more granular control and flexibility. Here’s how it compares:
Policy Types: While PSA focuses solely on security, Kyverno allows the creation of policies for validation, mutation, and generation of resources. This means you can modify or generate new resources, not just enforce security rules.
Customizability: Kyverno supports custom policies that can enforce your organization’s compliance requirements. You can write policies that govern specific resource types, such as ensuring that all deployments have certain labels or that container images come from a trusted registry.
Policy as Code: Kyverno policies are written in YAML, allowing for easy integration with CI/CD pipelines and GitOps workflows. This makes policy management declarative and version-controlled, which is not the case with PSA.
Audit and Reporting: With Kyverno, you can generate detailed audit logs and reports on policy violations, giving administrators a clear view of how policies are enforced and where violations occur. PSA lacks this built-in reporting capability.
Enforcement and Mutation: While PSA primarily enforces restrictions on pods, Kyverno allows not only validation of configurations but also modification of resources (mutation) when required. This adds an additional layer of flexibility, such as automatically adding annotations or labels.
When to Use Kyverno Over PSA
While PSA might be sufficient for simpler environments, Kyverno becomes a valuable tool in scenarios requiring:
Custom Compliance Rules: For example, enforcing that all containers use a specific base image or restricting specific container capabilities across different environments.
CI/CD Integrations: Kyverno can integrate into your CI/CD pipelines, ensuring that resources comply with organizational policies before they are deployed.
Complex Governance: When managing large clusters with multiple teams, Kyverno’s policy hierarchy and scope allow for finer control over who can deploy what and how resources are configured.
If your organization needs a more robust and flexible security solution, Kyverno is a better fit compared to PSA’s more generic approach.
Installing Kyverno
To start using Kyverno, you’ll need to install it in your Kubernetes cluster. This is a straightforward process using Helm, which makes it easy to manage and update.
After installation, Kyverno will begin enforcing policies across your cluster, but you’ll need to deploy some policies to get started.
Installing the Kyverno CLI
Separate from the in-cluster controller, Kyverno ships a CLI (kyverno) used to test and validate policies before they reach a cluster. It is the piece that makes policies reviewable in a pull request rather than discovered in production.
# macOS / Linux (Homebrew)
brew install kyverno
# Krew (as a kubectl plugin)
kubectl krew install kyverno
# Direct binary
curl -LO https://github.com/kyverno/kyverno/releases/latest/download/kyverno-cli_linux_x86_64.tar.gz
tar -xvf kyverno-cli_linux_x86_64.tar.gz && sudo mv kyverno /usr/local/bin/
The two commands worth knowing immediately:
# Apply a policy against manifests without touching a cluster
kyverno apply ./policies/ --resource ./manifests/
# Run the policy test suite defined in kyverno-test.yaml
kyverno test ./policies/
kyverno apply answers “would this policy have blocked this manifest?” and prints a policy report; kyverno test runs declarative test cases so policies have regression coverage like any other code. Wiring both into CI is covered step by step in integrating the Kyverno CLI into CI/CD pipelines with GitHub Actions.
The Kyverno Policy Library
Before writing anything yourself, check the official Kyverno policy library at kyverno.io/policies. It carries several hundred ready-made policies grouped by category — Pod Security Standards equivalents, best-practice rules, multi-tenancy controls, image verification, and cleanup policies. Most teams find that their first ten policies already exist there, and the realistic starting point is to adopt the Pod Security Standards set in audit mode, see what your cluster is already violating, and only then start writing custom rules.
Deploying Policies with Kyverno
Kyverno policies are written in YAML, just like Kubernetes resources, which makes them easy to read and manage. You can find several ready-to-use policies from the Kyverno Policy Library, or create your own to match your requirements.
Here is an example of a simple validation policy that ensures all pods use trusted container images from a specific registry:
This policy will automatically block the deployment of any pod that uses an image from a registry other than myregistry.com.
Applying the Policy
To apply the above policy, save it to a YAML file (e.g., trusted-registry-policy.yaml) and run the following command:
kubectl apply -f trusted-registry-policy.yaml
Once applied, Kyverno will enforce this policy across your cluster.
Viewing Kyverno Policy Reports
Kyverno generates detailed reports on policy violations, which are useful for audits and tracking policy compliance. To check the reports, you can use the following commands:
List all Kyverno policy reports:
kubectl get clusterpolicyreport
Describe a specific policy report to get more details:
These reports can be integrated into your monitoring tools to trigger alerts when critical violations occur.
Frequently Asked Questions
What is Kyverno used for?
Kyverno is a Kubernetes policy engine used to validate, mutate and generate cluster resources. Typical uses are enforcing that pods declare resource limits, blocking images from untrusted registries, requiring labels on namespaces, automatically generating a NetworkPolicy or ConfigMap in every new namespace, and verifying image signatures before a workload is admitted. Policies are written as YAML and applied as Kubernetes objects.
Kyverno vs OPA Gatekeeper: which should I use?
Use Kyverno if your policies only target Kubernetes and you want them written in YAML with no new language to learn — it also handles mutation and resource generation natively. Use OPA Gatekeeper if you need highly expressive or programmatic policy logic, or if you want to reuse the same Rego policies outside Kubernetes (APIs, microservices, Terraform). Both are CNCF Graduated and have converged on feature parity for core admission control, so the decision is usually about language and scope rather than capability.
Is Kyverno free? Who is the company behind it?
The Kyverno project is free and open source under the Apache-2.0 licence, and it is a CNCF Graduated project, so it is not controlled by a single vendor. It was created by Nirmata, who remain the largest contributor and offer commercial support and long-term support. Several other vendors also provide commercial Kyverno support, including Giant Swarm, InfraCloud, BlakYaks and Kodekloud.
How do I install the Kyverno CLI?
The quickest routes are brew install kyverno on macOS or Linux, or kubectl krew install kyverno to use it as a kubectl plugin. Binaries for each platform are also published on the GitHub releases page. The CLI is separate from the in-cluster controller: it is used to run kyverno apply and kyverno test against manifests in CI, before policies ever reach a cluster.
Do I still need Kyverno now that Kubernetes has ValidatingAdmissionPolicy?
For simple rules, often not — ValidatingAdmissionPolicy with CEL is built into Kubernetes, needs no extra components, and should be your first choice when the rule can be expressed in CEL. But CEL cannot make external calls or maintain state, and it only validates. Image-signature verification, cross-resource lookups, mutating resources on admission and generating new resources all still require Kyverno or Gatekeeper. Many clusters run both.
What is the difference between Kyverno and Pod Security Admission?
Pod Security Admission is built into Kubernetes and enforces the three fixed Pod Security Standards profiles (privileged, baseline, restricted) at namespace level. It is simple but not extensible — you cannot add a rule of your own. Kyverno is a general policy engine: it can reproduce the PSS profiles and then go further with custom rules, mutation and generation, and it applies to any resource type rather than pods alone.
Conclusion
Kyverno offers a flexible and powerful way to enforce policies in Kubernetes, making it an essential tool for organizations that need more than the basic capabilities provided by PSA. Whether you need to ensure compliance with internal security standards, automate resource modifications, or integrate policies into CI/CD pipelines, Kyverno’s extensive feature set makes it a go-to choice for Kubernetes governance.
For now, start with the out-of-the-box policies available in Kyverno’s library. In future posts, we’ll dive deeper into creating custom policies tailored to your specific needs.
In Kubernetes, security is a key concern, especially as containers and microservices grow in complexity. One of the essential features of Kubernetes for policy enforcement is Pod Security Admission (PSA), which replaces the deprecated Pod Security Policies (PSP). PSA provides a more straightforward and flexible approach to enforce security policies, helping administrators safeguard clusters by ensuring that only compliant pods are allowed to run.
This article will guide you through PSA, the available Pod Security Standards, how to configure them, and how to apply security policies to specific namespaces using labels.
What is Pod Security Admission (PSA)?
PSA is a built-in admission controller introduced in Kubernetes 1.23 to replace Pod Security Policies (PSPs). PSPs had a steep learning curve and could become cumbersome when scaling security policies across various environments. PSA simplifies this process by applying Kubernetes Pod Security Standards based on predefined security levels without needing custom logic for each policy.
With PSA, cluster administrators can restrict the permissions of pods by using labels that correspond to specific Pod Security Standards. PSA operates at the namespace level, enabling better granularity in controlling security policies for different workloads.
Pod Security Standards
Kubernetes provides three key Pod Security Standards in the PSA framework:
Privileged: No restrictions; permits all features and is the least restrictive mode. This is not recommended for production workloads but can be used in controlled environments or for workloads requiring elevated permissions.
Baseline: Provides a good balance between usability and security, restricting the most dangerous aspects of pod privileges while allowing common configurations. It is suitable for most applications that don’t need special permissions.
Restricted: The most stringent level of security. This level is intended for workloads that require the highest level of isolation and control, such as multi-tenant clusters or workloads exposed to the internet.
Each standard includes specific rules to limit pod privileges, such as disallowing privileged containers, restricting access to the host network, and preventing changes to certain security contexts.
Setting Up Pod Security Admission (PSA)
To enable PSA, you need to label your namespaces based on the security level you want to enforce. The label format is as follows:
This setup enforces the baseline standard while issuing warnings and logging violations for restricted-level rules.
Example: Configuring Pod Security in a Namespace
Let’s walk through an example of configuring baseline security for the dev namespace. First, you need to apply the PSA labels:
kubectl create namespace dev
kubectl label --overwrite ns dev pod-security.kubernetes.io/enforce=baseline
Now, any pod deployed in the dev namespace will be checked against the baseline security standard. If a pod violates the baseline policy (for instance, by attempting to run a privileged container), it will be blocked from starting.
You can also combine warn and audit modes to track violations without blocking pods:
kubectl label --overwrite ns dev pod-security.kubernetes.io/enforce=baseline pod-security.kubernetes.io/warn=restricted pod-security.kubernetes.io/audit=privileged
In this case, PSA will allow pods to run if they meet the baseline policy, but it will issue warnings for restricted-level violations and log any privileged-level violations.
Applying Policies by Default
One of the strengths of PSA is its simplicity in applying policies at the namespace level, but administrators might wonder if there’s a way to apply a default policy across new namespaces automatically. As of now, Kubernetes does not natively provide an option to apply PSA policies globally by default. However, you can use admission webhooks or automation tools such as OPA Gatekeeper or Kyverno to enforce default policies for new namespaces.
Conclusion
Pod Security Admission (PSA) simplifies policy enforcement in Kubernetes clusters, making it easier to ensure compliance with security standards across different environments. By configuring Pod Security Standards at the namespace level and using labels, administrators can control the security level of workloads with ease. The flexibility of PSA allows for efficient security management without the complexity associated with the older Pod Security Policies (PSPs).
Managing Kubernetes resources effectively can sometimes feel overwhelming, but Helm, the Kubernetes package manager, offers several commands and flags that make the process smoother and more intuitive. In this article, we’ll dive into some lesser-known Helm commands and flags, explaining their uses, benefits, and practical examples.
These advanced commands are essential for mastering Helm in production. For the complete toolkit including fundamentals, testing, and deployment patterns, visit our Helm package management guide.
1. helm get values: Retrieving Deployed Chart Values
The helm get values command is essential when you need to see the configuration values of a deployed Helm chart. This is particularly useful when you have a chart deployed but lack access to its original configuration file. With this command, you can achieve an “Infrastructure as Code” approach by capturing the current state of your deployment.
Usage:
helm get values <release-name> [flags]
Example:
To get the values of a deployed chart named my-release:
helm get values my-release --namespace my-namespace
This command outputs the current values used for the deployment, which is valuable for documentation, replicating the environment, or modifying deployments.
2. Understanding helm upgrade Flags: --reset-values, --reuse-values, and --reset-then-reuse
The helm upgrade command is typically used to upgrade or modify an existing Helm release. However, the behavior of this command can be finely tuned using several flags: --reset-values, --reuse-values, and --reset-then-reuse.
--reset-values: Ignores the previous values and uses only the values provided in the current command. Use this flag when you want to override the existing configuration entirely.
Example Scenario: You are deploying a new version of your application, and you want to ensure that no old values are retained.
--reuse-values: Reuses the previous release’s values and merges them with any new values provided. This flag is useful when you want to keep most of the old configuration but apply a few tweaks.
Example Scenario: You need to add a new environment variable to an existing deployment without affecting the other settings.
--reset-then-reuse: A combination of the two. It resets to the original values and then merges the old values back, allowing you to start with a clean slate while retaining specific configurations.
Example Scenario: Useful in complex environments where you want to ensure the chart is using the original default settings but retain some custom values.
3. helm lint: Ensuring Chart Quality in CI/CD Pipelines
The helm lint command checks Helm charts for syntax errors, best practices, and other potential issues. This is especially useful when integrating Helm into a CI/CD pipeline, as it ensures your charts are reliable and adhere to best practices before deployment.
Usage:
helm lint <chart-path> [flags]
<chart-path>: Path to the Helm chart you want to validate.
Example:
helm lint ./my-chart/
This command scans the my-chart directory for issues like missing fields, incorrect YAML structure, or deprecated usage. If you’re automating deployments, integrating helm lint into your pipeline helps catch problems early. By adding this command in your CICD pipeline, you ensure that any syntax or structural issues are caught before proceeding to build or deployment stages. You can lear more about helm testing in the linked article
4. helm rollback: Reverting to a Previous Release
The helm rollback command allows you to revert a release to a previous version. This can be incredibly useful in case of a failed upgrade or deployment, as it provides a way to quickly restore a known good state.
Usage:
helm rollback <release-name> [revision] [flags]
[revision]: The revision number to which you want to roll back. If omitted, Helm will roll back to the previous release by default.
Example:
To roll back a release named my-release to its previous version:
helm rollback my-release
To roll back to a specific revision, say revision 3:
helm rollback my-release 3
This command can be a lifesaver when a recent change breaks your application, allowing you to quickly restore service continuity while investigating the issue.
5. helm verify: Validating a Chart Before Use
The helm verify command checks the integrity and validity of a chart before it is deployed. This command ensures that the chart’s package file has not been tampered with or corrupted. It’s particularly useful when you are pulling charts from external repositories or using charts shared across multiple teams.
Usage:
helm verify <chart-path>
Example:
To verify a downloaded chart named my-chart:
helm verify ./my-chart.tgz
If the chart passes the verification, Helm will output a success message. If it fails, you’ll see details of the issues, which could range from missing files to checksum mismatches.
Conclusion
Leveraging these advanced Helm commands and flags can significantly enhance your Kubernetes management capabilities. Whether you are retrieving existing deployment configurations, fine-tuning your Helm upgrades, or ensuring the quality of your charts in a CI/CD pipeline, these tricks help you maintain a robust and efficient Kubernetes environment.
Five More Flags Worth Knowing
helm template --show-only templates/deployment.yaml renders a single template instead of the whole chart — the fastest way to debug one manifest without scrolling through hundreds of lines.
helm get manifest myrelease --revision 3 shows exactly what revision 3 deployed. Combined with diff <(helm get manifest r --revision 3) <(helm get manifest r --revision 4) you get a precise answer to “what changed between these two upgrades”.
helm status myrelease -o json makes release state scriptable — jq -r .info.status in CI gates is cleaner than parsing human output.
helm lint --strict turns warnings into errors. Without --strict, lint passes charts that will annoy you later; in CI there is no reason not to use it.
--post-renderer pipes Helm’s rendered output through any executable before applying — the standard escape hatch for applying kustomize patches to third-party charts you do not control, without forking them.
And one plugin that saves upgrades: mapkubeapis (helm plugin install https://github.com/helm/helm-mapkubeapis) rewrites release metadata that references Kubernetes APIs removed in newer versions — the fix for upgrades failing with “unable to build kubernetes objects” after a cluster upgrade.
Frequently Asked Questions
How do I see what a Helm release will change before applying it?
Use helm diff upgrade (from the helm-diff plugin) to see a rendered diff against the live release, or helm upgrade --dry-run --debug for the rendered manifests. For a diff against what is actually running in the cluster, helm get manifest piped to kubectl diff -f - works without plugins.
What does helm upgrade –reuse-values actually do?
It merges your new --set/-f overrides on top of the values stored from the previous release, ignoring any changes in the chart’s default values.yaml. That last part surprises people: after a chart version bump, new defaults do not apply. Prefer --reset-then-reuse-values or explicit values files in CI.
How can I roll back a Helm release to a specific revision?
helm history <release> lists revisions; helm rollback <release> <revision> restores one. The rollback itself creates a new revision, so the history is never rewritten — and --cleanup-on-fail removes resources created by the failed upgrade.
How do I get the values a deployed release was installed with?
helm get values <release> shows the user-supplied values; add --all to include chart defaults. Combined with helm get manifest, that reconstructs exactly what was deployed and why.
Istio has become an essential tool for managing HTTP traffic within Kubernetes clusters, offering advanced features such as Canary Deployments, mTLS, and end-to-end visibility. However, some tasks, like exposing a TCP port using the Istio IngressGateway, can be challenging if you’ve never done it before. This article will guide you through the process of exposing TCP ports with Istio Ingress Gateway, complete with real-world examples and practical use cases.
Understanding the Context
Istio is often used to manage HTTP traffic in Kubernetes, providing powerful capabilities such as traffic management, security, and observability. The Istio IngressGateway serves as the entry point for external traffic into the Kubernetes cluster, typically handling HTTP and HTTPS traffic. However, Istio also supports TCP traffic, which is necessary for use cases like exposing databases or other non-HTTP services running in the cluster to external consumers.
Exposing a TCP port through Istio involves configuring the IngressGateway to handle TCP traffic and route it to the appropriate service. This setup is particularly useful in scenarios where you need to expose services like TIBCO EMS or Kubernetes-based databases to other internal or external applications.
Steps to Expose a TCP Port with Istio IngressGateway
1.- Modify the Istio IngressGateway Service:
Before configuring the Gateway, you must ensure that the Istio IngressGateway service is configured to listen on the new TCP port. This step is crucial if you’re using a NodePort service, as this port needs to be opened on the Load Balancer.
After applying these configurations, the Istio IngressGateway will expose the TCP port to external traffic.
Practical Use Cases
Exposing TIBCO EMS Server: One common scenario is exposing a TIBCO EMS (Enterprise Message Service) server running within a Kubernetes cluster to other internal applications or external consumers. By configuring the Istio IngressGateway to handle TCP traffic, you can securely expose EMS’s TCP port, allowing it to communicate with services outside the Kubernetes environment.
Exposing Databases: Another use case is exposing a database running within Kubernetes to external services or different clusters. By exposing the database’s TCP port through the Istio IngressGateway, you enable other applications to interact with it, regardless of their location.
Exposing a Custom TCP-Based Service: Suppose you have a custom application running within Kubernetes that communicates over TCP, such as a game server or a custom TCP-based API service. You can use the Istio IngressGateway to expose this service to external users, making it accessible from outside the cluster.
Conclusion
Exposing TCP ports using the Istio IngressGateway can be a powerful technique for managing non-HTTP traffic in your Kubernetes cluster. With the steps outlined in this article, you can confidently expose services like TIBCO EMS, databases, or custom TCP-based applications to external consumers, enhancing the flexibility and connectivity of your applications.