Policy as code for Terraform: the gate is easy, the pipeline is not
Most teams that run Terraform have had this meeting.
Someone opened a pull request. The plan said 3 to destroy. Four hundred lines of diff scrolled past. Two people approved it. A queue that another service had been writing to for two years stopped existing.
Nobody was careless. The warning was on the screen. It was on line 312, and people do not read line 312.
Asking reviewers to try harder does not fix this. What fixes it is a program that reads the plan, spots the dangerous change, and stops the merge until a human says in writing that they meant it.
We built that gate for about twenty Terraform root modules. This article explains how the mechanism works, and what went wrong while we set it up. The examples are simplified and the infrastructure is generic. The mistakes are real.
One thing to clear up first, because people arrive at this problem holding the wrong tool. OPA Gatekeeper will not help you: it is a Kubernetes admission controller, it lives in the cluster’s API server, and it cannot read HCL or a Terraform plan. What it shares with the setup below is Rego, the policy language, which is worth something if your cluster policies are already written in it. The tool that reads a Terraform plan is Conftest, a small wrapper around OPA that takes a structured file, runs Rego against it, and returns an exit code.
How the gate works
This part is short, and that is the point. The policy engine took two days. Everything after it took weeks.
Check the plan, not the code. Conftest can read HCL directly, which needs no credentials and covers the whole repository at once. It is nearly useless for the rules that matter, because source code does not know what exists today. You delete a resource by removing its block, and a policy cannot match on an absence. State has the opposite problem: it describes what exists, not what is about to change. Only the plan carries the intent, with values resolved:
terraform plan -out=tfplan
terraform show -json tfplan > tfplan.json
That choice sets the boundary of the whole system, so say it out loud before anyone assumes otherwise.
A policy only protects what is planned.
A module with no plan in CI has no rules applied to it, however many rules you have written.
What the plan file says. Under resource_changes there is one entry per resource, and change.actions carries the decision:
{
"address": "aws_s3_bucket.exports",
"type": "aws_s3_bucket",
"change": {
"actions": ["delete"],
"before": { "bucket": "...", "force_destroy": false },
"after": null
}
}
The vocabulary is small: ["no-op"] for most of the file, ["read"] for data sources, the obvious ["create"], ["update"] and ["delete"], and then the two that catch people out. ["delete", "create"] is a replacement. ["create", "delete"] is the same thing where the new resource is built first.
A replacement destroys the old resource. For anything holding data that is the same loss as a deletion, but it reads differently in a diff and gets described in pull requests as a rename or a move. So test for membership:
"delete" in resource.change.actions
not equality against ["delete"]. We found this because someone asked what happens when a key changes. In our case one system keys its resources by strings in a map, some of them old identifiers that look untidy. Cleaning one up would have destroyed years of history behind a diff that looks cosmetic.
before and after hold resolved values, which gives you the second kind of rule: not “is this being destroyed” but “is this setting moving the wrong way”. A flag that takes a record out from behind a proxy. Deletion protection switched off. A retention period shortened. One word each, and none of them look urgent.
A rule, in full. They are all about this size.
package main
import rego.v1
stateful_types := {
"aws_dynamodb_table",
"aws_s3_bucket",
"aws_sqs_queue",
}
destroyed_stateful contains resource if {
some resource in input.resource_changes
resource.type in stateful_types
"delete" in resource.change.actions
}
deny contains msg if {
not destroy_allowed
some resource in destroyed_stateful
msg := sprintf(
"%s (%s) will be destroyed by this plan. Add the allow-destroy label to the pull request if that is intended.",
[resource.address, concat(", ", resource.change.actions)],
)
}
Conftest binds the plan to input. deny is a set, not a function: Rego tries every match and collects every message, so ten bad resources produce ten messages in one pass and no loop is written. An empty set is what passing looks like from the inside.
If that syntax looks unfamiliar, you have read older material. OPA 1.0 made if and contains mandatory, so the pre-1.0 form still shown in most tutorials, deny[msg] { ... } without either keyword, no longer compiles without the --v0-compatible flag. Iteration with [_] still works; some ... in just reads better.
The sentence inside sprintf is the part anyone will actually read. “Policy violation: rule 7 failed” is an obstacle. A line that names the resource, says what is lost and gives the next step is help. We spend more time on the message than on the logic.
The escape hatch. Conftest also has warn, which reports without failing. For a week of watching that is fine; as a permanent setting it is a trap, because a warning on every pull request stops being read within a month.
Deny by default instead, with a deliberate way to say yes. Ours is a label on the pull request, passed to the policy as data:
destroy_allowed if {
some label in data.pr.labels
label == "allow-destroy"
}
With no label the list is empty, destroy_allowed is undefined, and the rule denies. The policy does not change between a blocked run and an approved one, which is an easy thing to explain to an auditor. The label also stays visible in the pull request list and in the history, so “when did we drop that table, and who agreed” is a search rather than an excavation.
Use separate labels for separate decisions. We have three: one for destroying a resource, one for taking a DNS record out from behind the proxy, one for changing quality thresholds that apply to every project. Someone who thought about the first has not necessarily thought about the second, and a single allow-everything label turns three decisions into one habit.
Five things that went wrong
None of these are about Rego. Every one of them showed up in a real run, after the rules were written and working, and each took longer to find than the rule it broke.
Here is the whole pipeline, with the numbers marking where each one lives.
pull request opened, updated or labelled
│
▼
┌─────────────────────────────────┐ ① the job holds credentials, and runs
│ terraform init && plan -out │ configuration written in the branch
└────────────────┬────────────────┘
▼
┌─────────────────────────────────┐
│ terraform show -json tfplan │──▶ tfplan.json sensitive values in clear
│ terraform show tfplan │──▶ plan.txt sensitive values redacted
└────────────────┬────────────────┘
▼
┌─────────────────────────────────┐ ③ the override labels are read here,
│ conftest test tfplan.json │ and the payload is the wrong source
│ --data pr_data.json │
└────────────────┬────────────────┘
▼
┌─────────────────────────────────┐ ④ comment 65,536 chars, summary 1 MiB
│ comment · run summary · artifact│ ② the upload is skipped after a failure
└────────────────┬────────────────┘
▼
status check on the pull request
Number five is not on the diagram, because it is about which modules reach it at all.
1. The credentials, and what a plan can do with them
A plan needs to read state, and usually the infrastructure too. Our first version used a long-lived access key stored as a repository secret.
Replacing that with OIDC federation was the biggest single improvement. The pipeline gets a short-lived token from the CI provider. The cloud provider checks its signature and its claims, then hands back credentials that live for one job. There is no stored key to leak.
The trust condition is the whole security boundary, so make it exact. Matching the identity claim with a wildcard across an organization lets every repository in that organization assume the role. Match the full claim. If you want an approval step, use the CI provider’s environment claim, because anything written only in the workflow file can be edited by whoever opened the pull request.
In practice that is two pieces. The workflow asks for a token and trades it:
permissions:
contents: read
id-token: write # without this no token is minted
steps:
- uses: aws-actions/configure-aws-credentials@v6
with:
role-to-assume: arn:aws:iam::<account>:role/terraform-plan
aws-region: us-east-1
And the role decides who may do that:
{
"Effect": "Allow",
"Action": "sts:AssumeRoleWithWebIdentity",
"Principal": { "Federated": "arn:aws:iam::<account>:oidc-provider/token.actions.githubusercontent.com" },
"Condition": {
"StringEquals": {
"token.actions.githubusercontent.com:aud": "sts.amazonaws.com",
"token.actions.githubusercontent.com:sub": "repo:<org>/<repo>:pull_request"
}
}
}
Note StringEquals on the whole sub, not StringLike with a wildcard. repo:<org>/* would let any repository in the organization assume this role. If you want a human approval in front of a sensitive module, gate the job on a CI environment and match ...:environment:<name> instead: that claim only appears when the job really declares the environment, so removing the declaration from the workflow file does not get you the role.
Now the uncomfortable part, and the one your security reviewer will find:
terraform planruns configuration from the branch under review.
Someone can add a data source that reads the job’s environment and sends it somewhere. Every credential in that job is reachable by anyone who can open a pull request, and so is every state file the job’s role can read.
Most CI systems do not give secrets to pull requests from forks, so in practice the exposure is to people who already have write access. That is not nothing: “already trusted” is not the same as “cannot be phished”. Give the role read-only access, scope it to the state it needs, and decide before you add your most sensitive module whether its plan should sit behind an approval.
2. A condition that skipped the step we needed most
Our pipeline uploads the readable plan as an artifact. The step had a condition: only run if the plan succeeded.
In GitHub Actions, a condition that does not mention a status function gets an implicit success() added to it. So as soon as the policy check failed, the upload was skipped. Meanwhile the pull request comment still pointed at an artifact that did not exist, exactly when a reviewer most needed to read the plan.
# skipped as soon as any earlier step fails, artifact and all
- name: Upload plan
if: steps.plan.outcome == 'success'
# runs on its own terms
- name: Upload plan
if: always() && steps.plan.outcome == 'success'
The fix is one word, always(). The lesson is bigger: after a step fails, every condition below it means something different.
3. Labels the pipeline could not see
The override label was read from the event that started the run.
A reviewer’s natural sequence is: check fails, read the message, add the label, re-run the failed job. That does not work. Re-running replays the original event, labels included. The new label is invisible and the check fails again in exactly the same way. From the outside this looks like “the override is broken”.
Two changes fixed it. The job now asks the API for the pull request’s current labels instead of trusting the stored event. And the workflow also triggers when a label is added or removed, so applying the label starts a fresh run by itself.
on:
pull_request:
# the default is [opened, synchronize, reopened]: labelling fires nothing
types: [opened, synchronize, reopened, labeled, unlabeled]
# not github.event.pull_request.labels: that is the payload the run started with
gh pr view "$PR_NUMBER" --repo "$REPO" \
--json labels --jq '{pr: {labels: [.labels[].name]}}' > pr_data.json
echo "Labels seen by the policies: $(jq -c '.pr.labels' pr_data.json)"
conftest test tfplan.json --policy policy --data pr_data.json
The job also prints the labels it saw. That turns an argument into a five-second check.
4. Output that outgrew the comment
Our first version pasted the whole plan into a pull request comment and cut it off when it got too long. Fine for a small environment. Not fine for a real one: the comment limit is 65,536 characters, and reviewers got a plan that stopped mid-resource.
What works is three layers.
The comment holds the summary line and one line per changed resource, built from the plan JSON instead of the text output:
jq -r '.resource_changes[]
| select(.change.actions != ["no-op"] and .change.actions != ["read"])
| "\(.change.actions | join("+")) \(.address)"' tfplan.json | sort
which gives one line per real change:
delete aws_s3_bucket.exports
delete+create aws_dynamodb_table.sessions
update aws_sqs_queue.jobs
Five hundred resources is about 25 KB, which fits easily.
The full readable plan goes into the pipeline’s run summary page, which renders in the browser with a 1 MiB budget and needs no download.
# readable plan, no refresh noise, safe to publish
terraform show -no-color tfplan > plan.txt
{
echo "## Plan for ${DIR}"
echo '```hcl'
head -c 900000 plan.txt
echo '```'
} >> "$GITHUB_STEP_SUMMARY"
The same file is attached as an artifact for anyone who wants to search it locally.
One detail worth flagging to your security and legal reviewers: attach the human-readable plan, not the JSON. Terraform hides sensitive attributes in its rendered output. The JSON contains them. An artifact can be downloaded by everyone with read access and stays for its retention period, which by default is far longer than any review needs. Shorten that, and remember each run leaves its own copy instead of replacing the last one.
5. Coverage that looked better than it was
The pull request comment was green long before the rules protected anything important, because most modules produced no plan at all.
If you take one operational point from this article, take this one. Track which modules are actually planned in CI. Publish the list. Treat it as the real coverage number.
Ours is two commands and a diff. Every root module declares a backend, so the first list is free:
# every root module in the repository
grep -rl 'backend "s3"' --include='*.tf' . | xargs -n1 dirname | sort -u
# the ones the pipeline actually plans
yq '.on.pull_request.paths[]' .github/workflows/terraform-plan.yaml
The gap between those two lists is the honest answer to “is production covered”. Ours started at two modules out of twenty. Nobody would have guessed that from the green checkmarks.
Rules are cheap. Plans are the expensive part, because each new module needs credentials, network access, and sometimes an approval design.
Test the rules, then test the plumbing
Policy code is code, and it has an unpleasant failure mode: a broken rule does not crash. It just stops denying, and everything turns green. The thing you have to design against is silence.
That needs two layers.
Unit tests, in Rego. Conftest runs files named *_test.rego. A test hands the policy a small hand-written document instead of a real plan:
replace_plan := {"resource_changes": [{
"address": "aws_s3_bucket.media",
"type": "aws_s3_bucket",
"change": {"actions": ["delete", "create"]},
}]}
test_replacement_is_denied if {
count(deny) == 1 with input as replace_plan
}
test_label_overrides_deny if {
count(deny) == 0 with input as replace_plan
with data.pr.labels as ["allow-destroy"]
}
No Terraform runs. No credentials. These finish in about a second, and they cover the logic: the right resource types, replacement treated as deletion, the override working, the wrong override not working.
One test of the command itself. Unit tests run inside the policy engine, so they cannot see how it gets called. In our pipeline it is called with a policy directory, a data file, and an exit code that a shell script reads. All three can break without touching a single rule.
So we keep one hand-written plan that breaks every rule at once. The pipeline checks two things:
- The real command rejects it.
- The same command with every override label accepts it.
The first check is inverted: if the command passes, the step fails, because this file is supposed to be rejected. That inversion caught a folder rename that had quietly stopped half our rules from loading.
conftest verify --policy policy
echo '{"pr": {"labels": []}}' > no_labels.json
echo '{"pr": {"labels": ["allow-destroy", "allow-unproxied", "allow-threshold-change"]}}' > all_labels.json
# inverted on purpose: the fixture breaks every rule, so passing is the failure
if conftest test tests/plan_violations.json \
--policy policy --data no_labels.json; then
echo "The fixture violates every rule, but conftest passed it."
exit 1
fi
conftest test tests/plan_violations.json \
--policy policy --data all_labels.json
Print a line before it saying the failures below are expected. Otherwise the next person to open that log sees six red FAIL lines in a green job and files a bug.
Two practical notes. Keep that file outside the policy folder, so nothing can mistake a test document for a policy input. And make it a rule that every new policy adds its own violation to the file, otherwise the file slowly stops covering everything and the check goes green for the wrong reason.
Before this we did something else: when a policy changed, the pipeline ran a real plan against one environment as a smoke test. It worked, and everyone disliked it, because a pull request that touched only policy files would start planning infrastructure it had never touched. The fixture replaced it. It is faster, needs no credentials, and tests more.
What this does not do
Being clear about the limits is what keeps the gate trusted.
It does not see changes made outside Terraform. Someone deleting a resource in a web console is invisible here.
It does not cover modules that produce no plan.
It is not a security scanner. The established scanners cover unencrypted storage, public buckets and open security groups better than anything you will write by hand, and they need no policy authoring at all. Use them for known bad configuration. Use custom policy for facts only your organization knows: which resources cannot be recreated, which identifiers other systems depend on, which thresholds were chosen deliberately.
It does not stop a determined person with merge rights. It makes an accident expensive and a deliberate act recorded.
And it does not survive neglect. Write twenty rules in a planning meeting and half of them will be dead within a quarter, with everyone adding the override label without reading the message. Ours grow one at a time. Each one traces back to an incident, a near miss, or a review comment where someone asked what stops this from happening again. Nine rules across four providers. Each has a story.
If you are starting on Monday
Write the tests before the second rule, not after the tenth.
Keep one fixture that breaks everything, and keep it up to date.
Start with a single rule about destruction. Everyone agrees with it, so nobody argues in review.
Make the message a sentence you would be glad to read at the moment it stops you.
Put the override behind something that sticks around and says what was approved.
Federate your credentials before you widen what the pipeline can reach, not after.
And write down which modules are covered, so that when someone asks whether the gate protects production, the answer is a list and not a hope.