Adopting OPA Gatekeeper: mechanics and failure modes
Why admission control
Every cluster runs on rules that exist only in people’s heads. Images come from our registry. Nothing runs privileged. Tags are pinned. Everyone agrees, and then a debugging session at two in the morning mounts the node’s filesystem into a pod, and the rule quietly stops being true.
Code review catches some of this. It cannot catch what never reaches review: a kubectl run from a laptop, a Helm chart whose values nobody read, a manifest applied by a script written three years ago by someone who has since left. The gap is not discipline. The rules are simply not enforced at the point where objects actually arrive.
Admission control closes that gap by moving the rules into the API server’s request path. Every create and update is offered to a policy engine before it is persisted. OPA Gatekeeper is one implementation: policies are written in Rego, shipped as ordinary Kubernetes resources, and evaluated both at admission time and continuously against what already exists.
This article covers how Gatekeeper works in the places that matter operationally, then the four failure modes encountered while rolling it out across three clusters over several weeks. The mechanics are documented upstream. The failure modes are the expensive part, and most of them are not.
Examples are anonymised. Registry names, namespaces and hostnames below are placeholders; the shapes of the problems are not.
The moving parts
Gatekeeper splits a policy into two objects: a template that defines the rule, and a constraint that decides where the rule applies. The separation is the whole design, and it is what makes a policy library reusable across clusters that share nothing else.
A ConstraintTemplate carries the logic and, as a side effect, mints a new CRD:
apiVersion: templates.gatekeeper.sh/v1
kind: ConstraintTemplate
metadata:
name: k8sblocknodeport
spec:
crd:
spec:
names:
kind: K8sBlockNodePort # the CRD this creates
targets:
- target: admission.k8s.gatekeeper.sh
rego: |
package k8sblocknodeport
violation[{"msg": msg}] {
input.review.kind.kind == "Service"
input.review.object.spec.type == "NodePort"
msg := "Service of type NodePort is not allowed"
}
Applying that template creates a CRD named K8sBlockNodePort. A constraint is an instance of it, and instances carry scope and severity:
apiVersion: constraints.gatekeeper.sh/v1beta1
kind: K8sBlockNodePort
metadata:
name: block-nodeport-services
spec:
enforcementAction: deny
match:
kinds:
- apiGroups: [""]
kinds: ["Service"]
excludedNamespaces: [kube-system, ingress]
Five things about the Rego explain most of the surprises later. They are worth learning even if you never write a policy yourself.
violation is a set, not a yes/no. You never write “allow”. You list reasons to reject. An empty set means the object passes. Each element becomes one line in a report, so one bad container produces one violation and three produce three.
Lines in a rule are joined by AND. There is no operator for it. The rule fires only if every line holds.
Two rules with the same name are OR. That is how you write “or” in Rego: not with an operator, but by defining the rule twice. This matters more than it sounds, see failure mode 1.
A missing field usually means “allowed”. In Rego, reading a field that is not there gives undefined, not false. The rule stops and produces no violation. So a policy about spec.type ignores an object that has no spec.type, unless the author added a rule for that case.
input is what you are given. input.review.object is the object under review. input.parameters is whatever the constraint put in spec.parameters. Looping is written [_], and every match produces its own violation.
Two loops that do not talk to each other
Gatekeeper evaluates the same policies in two completely separate places, and conflating them is the single most common source of confusion.
kubectl / CI / controller
│
▼
┌──────────────┐ admission review ┌──────────────┐
│ API server │ ───────────────────▶ │ Gatekeeper │
│ │ ◀─────────────────── │ webhook │
└──────┬───────┘ allow / deny └──────────────┘
│
▼
( etcd )
┆ every N seconds
▼
┌──────────────┐ ┌───────────────────┐
│ Audit │ ┄┄┄┄┄┄▶ │ constraint status │
│ controller │ │ + metrics │
└──────────────┘ └───────────────────┘
The admission webhook is synchronous and sits in the request path. It sees an object once, at the moment someone tries to create or change it, and answers allow or deny. It never sees anything that already exists.
The audit controller is asynchronous. On an interval it walks objects already in the cluster, evaluates them against every constraint, and writes the results to each constraint’s status and to metrics. It blocks nothing, ever.
Three consequences follow, and all three bite in practice.
Zero violations does not mean nothing would be blocked. It means the audit found nothing. Objects created since the last audit run, and objects that were refused and so never created, are both missing from that number.
A new constraint reports zero until its first audit finishes. Zero-because-clean and zero-because-not-computed-yet look identical. Tell them apart with status.auditTimestamp: when every constraint shows the same recent timestamp, a full cycle has run and the numbers mean something.
Exemptions are not symmetric. The --exempt-namespace flag (controllerManager.exemptNamespaces in the Helm chart) silences the webhook only. The audit controller ignores it and walks every namespace anyway, as the upstream docs state. To exclude a namespace from both, either repeat the list in each constraint’s match.excludedNamespaces, or list it once in the cluster-wide Config resource with processes: ["*"]. The Config route still sends every request to the webhook, so a Gatekeeper outage still reaches that namespace; the flag does not. In review either one looks redundant next to the flag. It is not.
Enforcement, and what happens when the engine is down
Nothing is blocked until you add a constraint. Every restriction is opt-in, so the rollout can go one policy at a time.
Each constraint has an enforcementAction:
| Action | Admission | Audit | Use it for |
|---|---|---|---|
dryrun |
allows | records | measuring a policy before you commit to it |
warn |
allows, warns the author | records | telling people without breaking them |
deny |
rejects | records | enforcement |
One detail catches people out: if you leave the field out, the default is deny, not dryrun. A constraint written without it starts blocking the moment it lands. Always set it explicitly.
The next question is what happens when Gatekeeper itself is down. The webhook sits in front of every create in the cluster, so this matters.
The answer depends on failurePolicy:
Ignore: the API server carries on when the webhook does not answer. Policies stop being enforced during an outage. The cluster keeps working.Fail: the API server rejects the request. Policies hold. An engine outage becomes a cluster-wide write outage.
Most teams want Ignore, with a short webhook timeout. A policy engine that can take the cluster down with it is a bad trade when the policies are hygiene rules.
But be honest about what that means. Ignore makes the guarantee best-effort. Anyone who can disrupt the webhook can bypass the policy. If a policy is part of your security boundary, Fail is the right choice, and you handle the availability risk by running the engine properly, not by weakening the policy.
Either way, run the controllers with several replicas, a pod disruption budget and anti-affinity. Under Ignore an outage silently switches enforcement off, which is exactly when you want it on.
Policies match Pods; humans apply Deployments
Almost every interesting workload policy is about a pod spec, so the constraint matches Pod. Almost nobody applies a Pod. They apply a Deployment, a StatefulSet or a CronJob, and the pod appears later, created by a controller.
Match only Pod and the rejection lands in the wrong place. The Deployment is admitted happily; its ReplicaSet then fails to create pods, and the author sees a healthy-looking Deployment with zero ready replicas and an error buried in events.
ExpansionTemplate fixes this. It tells Gatekeeper how to derive the pod a workload would produce, so policies evaluate against that derived pod at the moment the workload is applied:
apiVersion: expansion.gatekeeper.sh/v1beta1
kind: ExpansionTemplate
metadata:
name: expand-workloads
spec:
applyTo:
- groups: ["apps"]
kinds: ["Deployment", "StatefulSet", "DaemonSet"]
versions: ["v1"]
templateSource: "spec.template"
generatedGVK:
kind: "Pod"
version: "v1"
With expansion in place, the rejection arrives on kubectl apply of the Deployment, and the message is prefixed to say it was implied by expansion. Jobs and CronJobs need their own templates, since the path to the pod spec differs.
Two details are easy to miss.
Expansion evaluates spec.template regardless of replica count. A Deployment scaled to zero still fails admission, because the pod spec is still being declared. Dormant workloads left at zero replicas are a real source of violations in clusters that keep a full service catalogue in every environment.
Deliberately leaving ReplicaSet out of the expansion chain avoids double-reporting. A Deployment and the ReplicaSet it owns would otherwise each produce the same violation.
Order policies by cost, not by severity
The instinct is to rank candidate policies by how bad the thing they prevent is, adopt the worst first, and work down. That ordering is wrong, and following it is how Gatekeeper rollouts stall.
What determines whether a policy ever reaches enforcement is not its value. It is how many objects already violate it. A policy that cannot be driven to zero stays in dryrun forever, and a dashboard full of permanent violations teaches everyone to ignore the dashboard, including for the policies that did reach zero.
So measure first. Render every overlay and chart the cluster deploys, evaluate the candidate policies against the rendered output offline, and count. The spread is usually enormous.
These are real counts from one production cluster’s rendered manifests, across roughly 50 workloads:
| Candidate policy | Violating containers |
|---|---|
| Block privileged containers | 0 |
| Block hostPID / hostIPC | 0 |
| Block hostNetwork / hostPort | 0 |
| Block hostPath volumes | 0 |
| Block NodePort / LoadBalancer services | 0 |
| Require resource requests | 6 |
| Require liveness and readiness probes | 5 |
| Require memory limits | 3 |
Require allowPrivilegeEscalation: false |
80 |
| Require read-only root filesystem | 80 |
| Require dropping all capabilities | 80 |
| Require running as non-root | 80 |
The top group is free. Nothing violates it, so it can go straight to deny on the day it is written, and it closes exactly the escape routes a container breakout would use.
The bottom group touches every workload in the estate, because no base manifest sets securityContext at all. Same security literature, same “best practice” list, and the cost runs from zero to every workload.
The ordering that works is value divided by violation count. Ship the free policies immediately. Schedule the expensive ones as projects with owners, and be honest that they are projects.
For anything not free, the ladder is: dryrun, measure, fix the causes, confirm a fresh full audit reports zero, then promote. Confirm means checking that every constraint shares one recent auditTimestamp, not glancing at a zero.
Failure mode 1: policies that fault absence
Two policies from the same upstream library, both described as pod hardening, differ by a single Rego rule, and that difference is the whole reason one costs nothing and the other costs a quarter.
The privileged-container policy flags only an explicit setting:
violation[{"msg": msg}] {
c := input_containers[_]
c.securityContext.privileged # only true when explicitly set
msg := sprintf("Privileged container is not allowed: %v", [c.name])
}
The privilege-escalation policy adds a second rule:
input_allow_privilege_escalation(c) {
not has_field(c, "securityContext") # no securityContext at all
}
input_allow_privilege_escalation(c) {
not c.securityContext.allowPrivilegeEscalation == false
}
Remember that several rules with the same name are OR. The first rule means a container with no securityContext block is a violation. That is deliberate upstream: the field defaults to permissive, so silence really is insecure. It is also why this policy flagged 80 containers where the privileged check flagged zero.
The lesson generalises past this one policy. Before adopting anything from a policy library, read the Rego and answer one question: does an absent field count as a violation?
If yes, the policy is not a rule you switch on. It is a migration across every workload you own, and it needs the schedule and the owner that implies. The same applies to read-only root filesystems, dropping all capabilities, running as non-root and seccomp profiles. Each of them flagged the entire estate for the same reason.
Failure mode 2: image matching is string matching
The registry allowlist policy compares the image string exactly as written. There is no normalisation, no registry resolution, no awareness that a bare name means Docker Hub. Given an allowlist containing docker.io/library/:
| Image as written | Result |
|---|---|
docker.io/library/postgres:18-alpine |
allowed |
postgres:18-alpine |
rejected |
registry.example.com/team/app:v1.2.3 |
allowed if that prefix is listed |
docker.io/team/app:v1.2.3 |
rejected unless that exact prefix is listed |
This is defensible. A bare name is an implicit Docker Hub pull, which is both a supply-chain exposure and a rate-limit exposure, and forcing it to be written out makes that explicit. But it produces two traps.
The first is the trailing slash. An allowlist entry of acme without the slash also permits acme-evil/backdoor:v1, because prefix matching does not know where a registry name ends. Every entry needs the trailing slash, and a test case should guard it. This is the kind of thing that is obvious once stated and invisible in review.
The second is asymmetry between the forms your team writes. If the allowlist contains myorg/ for your own images and docker.io/library/ for upstream ones, then myorg/app:v1 is correct and docker.io/myorg/app:v1 is rejected, while postgres:18 is rejected and docker.io/library/postgres:18 is correct. Both conventions are defensible; having both in one list is a trap that fires months later, when someone writes the fully-qualified form out of tidiness.
Write the rule down beside the allowlist, and test both spellings.
Failure mode 3: the objects nothing can see
This is the one that cost real money, and it is the reason to read the rest of this article.
A nightly job backed up the cluster’s databases. For each database it created a short-lived helper pod, ran a dump through it, streamed the result to object storage and deleted the pod. The pod was created imperatively, from a shell script stored in a ConfigMap:
IMAGE="postgres:18-alpine" # note: no registry prefix
kubectl run "${pod_name}" --image=${IMAGE} -n "${NAMESPACE}" ...
The registry allowlist went to deny. Every helper pod was rejected at admission from that moment on. Backups produced nothing for twelve days.
What makes this worth dwelling on is that every check we had was clean, and each was clean for a different and individually reasonable reason.
The rendered manifests were clean, because that pod is not in any manifest. It exists only as a string inside a shell script inside a ConfigMap. Rendering the overlay produces the ConfigMap, not the pod.
The policy unit tests were clean, for the same reason. They test manifests. There is no manifest.
The audit controller was clean, because audit only evaluates objects that exist. These pods were rejected at creation, so they never existed to be audited. A policy working perfectly and a policy matching nothing produce the same zero.
And the CronJob reported success throughout, which is the subject of the next section.
The general statement is uncomfortable: anything that creates pods at runtime is invisible to every static check in your pipeline. That includes backup and restore jobs, operators that spawn workers, CI runners that start build pods, kubectl debug, Helm hooks, and any tool shelling out to kubectl run.
There is no clever fix. There is a checklist item. Before promoting any policy to deny, grep the repository for imperative pod creation and evaluate each site by hand:
grep -rn 'kubectl run\|--image=' --include='*.sh' --include='*.yaml' .
In our case that search turned up six more sites beyond the one that broke, including the database restore script, which would have failed at the exact moment someone needed it most.
One more thing is worth checking at the same time. A rejection has to be received by something, and that something was written before anyone thought about admission control. Look at what the caller does with a failure: a loop that logs the error and carries on will turn a policy rejection into a silent no-op, and the run will still report success.
That is worth knowing before you promote a policy, not after. In our case the rejection was sent to /dev/null and the job exited zero, so the breakage stayed invisible far longer than it needed to.
Failure mode 4: two engines in one template
Kubernetes now has its own policy mechanism, ValidatingAdmissionPolicy. It runs CEL expressions inside the API server, with no webhook. Gatekeeper can compile constraints down to it, and many upstream templates now carry both implementations, a Rego block and a CEL block, in one ConstraintTemplate.
That causes two problems. One is cosmetic. One is not.
The cosmetic one: the two engines word their messages differently. Offline test tools run the Rego path. The cluster may run the CEL path. So a test that asserts on message text can pass locally and describe something the cluster never says. Assert that a violation happened. Do not assert its wording unless you know which engine answers.
The serious one is about failurePolicy. A native policy runs inside the API server. If it has failurePolicy: Fail and cannot be evaluated while the API server is starting, because the resources it needs are not loaded yet, then the API server starts rejecting requests it makes of itself. The node never finishes coming up.
We hit this on a control-plane node. Three properties make it nasty:
- Namespace exemptions do not help. The failing requests are not namespaced.
dryrundoes not help. The failure is in evaluating the policy, not in its verdict.- Recovery runs through the API server that is refusing to start.
This is a known upstream issue, not something exotic we did: see Gatekeeper #4530, “Control plane deadlock via Gatekeeper-generated ValidatingAdmissionPolicies”. The general webhook version of the same trap is described in the Kubernetes docs, which warn that a cluster can reach a state where self-healing is impossible.
Recovery means removing the policy objects out of band, through the control plane’s own configuration rather than kubectl. That is a bad thing to be learning at the time.
Kubernetes 1.36 added manifest-based admission control partly to close this bootstrap window. Until you are on it, treat failurePolicy: Fail on a native policy as a control-plane change, not a policy change. Roll it to one node. Restart that node on purpose. Only then continue.
Testing, and the edge of what testing covers
Gatekeeper ships gator, a CLI that embeds the OPA engine, so policies can be tested with no cluster at all. Two modes matter.
gator verify runs unit tests: a suite pairs a template with a constraint and a set of fixture objects, each with an expected outcome. gator test takes rendered manifests and a set of policies and reports what would be rejected.
A suite entry is unremarkable, which is the point:
- name: privileged-containers
template: ../templates/k8spspprivilegedcontainer.yaml
constraint: fixtures/constraint-deny.yaml
expansion: ../templates/expansion-workloads.yaml
cases:
- name: privileged-container-is-rejected
object: fixtures/pod-privileged.yaml
assertions:
- violations: yes
- name: absent-security-context-is-accepted
object: fixtures/pod-no-security-context.yaml
assertions:
- violations: no
That second case is worth copying. It pins the property from failure mode 1, that this particular policy tolerates a missing securityContext, so that a future template upgrade changing the behaviour shows up as a failing test rather than a broken deployment.
Two more habits proved their worth. Keep a negative control for prefix matching, a fixture like acme-evil/backdoor:v1 asserting that the trailing slash in the allowlist does its job. And run gator test in CI against rendered manifests with enforcement forced to deny, regardless of what the cluster is running, so a bad manifest fails a pull request rather than a deploy.
In a GitOps shop this is where enforcement actually lives. Objects reach the cluster from the repository, so CI is the gate that authors experience. Admission deny is the backstop for everything that does not come from the repository, which, as failure mode 3 showed, is the interesting part.
The boundary is worth stating plainly, because it is easy to over-trust a green pipeline. gator sees manifests. It does not see Helm releases you did not render, objects created by operators, or anything constructed at runtime. The gap between “CI is green” and “the cluster is clean” is exactly the set of objects nobody wrote down.
What changes the day you switch to deny
Enforcement changes how people work, and the changes are worth anticipating rather than discovering.
Debugging gets harder. kubectl debug node/<name> creates a pod with host namespaces and the node’s root filesystem mounted. Under a host-escape policy set it is rejected outright. kubectl debug on a pod defaults to a bare upstream image, which a registry allowlist rejects. Both are correct behaviour and both will be reported as “Gatekeeper is broken” the first time.
Decide the break-glass path in advance. A dedicated namespace excluded from the relevant constraints is cleaner than teaching people to flip policies to dryrun, because it is one known exception in one known place rather than a habit of disabling enforcement under pressure.
Most debugging does not need elevation. Worth knowing, because the instinct is to reach for host access. Kubernetes leaves CAP_NET_RAW in the default capability set, so tcpdump, ping, mtr and port scanning all work in an ordinary pod against the pod’s own network. Elevation is needed only for the node’s netfilter tables, other pods’ traffic, tracing node processes, and reading node files.
Scope your policies deliberately, and know the difference. A policy scoped by namespaces covers a listed set; one scoped by excludedNamespaces covers everything else, including namespaces created next month by someone who has not read your policy. The second is the safer default for host-escape rules and the riskier one for anything opinionated about images.
Keep the back-out one line. Setting a single constraint to dryrun should be a one-line change in one file, reviewable in seconds. Anything more elaborate will not be used at three in the morning.
Finally, watch the webhook itself. Constraint count, rejection rate by constraint, and webhook latency all belong on a dashboard. Rejections are the interesting signal: a spike means either an attack, or far more likely, that someone’s routine workflow just stopped working and they have not told you yet.
What it buys, and what it does not
Admission control is worth adopting. It is the only place a cluster rule becomes a fact rather than an agreement, and it costs one afternoon. The free tier of policies (no privileged containers, no host namespaces, no hostPath, no NodePort or LoadBalancer services, no arbitrary external IPs) closes the routes a container escape actually uses.
What it does not buy is the thing the marketing implies. It does not make a cluster compliant, and it does not replace the work of fixing what the policies find. The expensive policies stay expensive; a tool that tells you 80 workloads run as root has not made any of them stop.
The adoption sequence that worked:
- Install the engine with no constraints. Confirm the webhook is healthy and
failurePolicyis deliberate. - Render every manifest you deploy and measure candidate policies offline. Rank by violations, not by severity.
- Adopt the zero-violation set straight to
deny. Write unit tests, including a negative control per policy. - Grep for imperative pod creation and evaluate every site by hand. This is the step that gets skipped.
- Put everything else in
dryrun. Fix causes. Promote only after a full audit cycle reports zero. - Alert on outcomes, not on the engine. A backup that stops producing backups matters more than a constraint count.
One distinction is worth carrying away. Admission control makes a rule true at the boundary of the cluster. It does not make what sits behind that boundary correct, and it does not decide what happens when a rule fires.
That second part is the one teams underestimate. A policy engine will do exactly what you asked of it, correctly, from the first night. Everything downstream of a rejection will behave the way it already behaved. Enforcement is very good at turning an unwritten assumption into a fact, and it is no substitute for knowing how your own systems fail.
Sources
- Handling Constraint Violations: enforcement actions and the
denydefault - Exempting Namespaces: webhook exemption versus audit
- Gatekeeper #4530: control plane deadlock via generated ValidatingAdmissionPolicies
- ValidatingAdmissionPolicy: Kubernetes documentation
- Kubernetes v1.36: manifest-based admission control