Skip to content

About

Deriving a least-privilege CloudFormation deployment policy: deny-first vs admin-first, measured. No LLM — every grant traceable to an observed denial.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Deriving a least-privilege CloudFormation deployment policy

Two experiments in working out exactly which IAM permissions the order-pipeline stack needs in order to deploy, and a comparison of the two methods.

Result: results/deploy-policy.compact.json — every grant traceable to an observed denial and scoped to a specific ARN, with none falling back to *. Verified to create, roll back, and delete the full 23-resource stack without admin intervention.

The committed artifact is 15 statements / 3,552 bytes. A later from-zero rerun produced the same 44 actions in 17 statements / 5,560 bytes: the loop captured lambda:DeleteEventSourceMapping against a per-deploy UUID as well as the wildcard, and the smaller figure reflects hand-scrubbing that is not in the code. Treat 3,552 as a best case, not a reproducible output.

The two methods

Deny-first (tools/discover.py). Start the CloudFormation service role at ReadOnlyAccess, deploy, harvest every AccessDenied, grant exactly that action on exactly that ARN, repeat until the stack completes. Then induce a rollback, then delete. Create and delete each discover permissions the other does not; rollback, in two separate runs, discovered none of its own — see Known limitations.

Admin-first (tools/admin_discover.py). Give the role AdministratorAccess, deploy once, and read CloudTrail for what it actually called. This is what AWS productised as IAM Access Analyzer policy generation.

What the comparison found

Deny-first Admin-first
Cost ~20 iterations, hours 1 iteration, 292s
Mutating actions 44 41
Read actions 0 (masked by ReadOnlyAccess) 52
Actions scoped to a real ARN 44 17
Actions scoped to * 0 24
Deploys the stack yes no — 11 of 23

The action counts look comparable. They are not the interesting number.

Least privilege lives in the Resource field, and that is where the methods diverge. A denial always names the resource that was refused. A successful CloudTrail event usually has an empty resources array and no error text, so there is nothing to scope to and the derivation falls back to *. The admin-first policy is smaller because it is less specific, not tighter.

Making the admin-first policy work took 10 further deny-first passes plus two fixes it could not make itself — at which point it had become a deny-first run wearing admin-first's starting set.

Three things admin-first structurally cannot learn:

  1. iam:PassRole. Authorised inside the calling API, so CloudTrail never logs it as its own event (documented). It halted the converge run at 21/23 until granted.
  2. Correct action names. CloudTrail reports API names; for several S3 operations these differ from the authorising action (PutBucketEncryption -> s3:PutEncryptionConfiguration). The static scan correctly rejects the API name, so the permission is lost rather than merely inert.
  3. Failure and replacement paths. lambda:RemovePermission, iam:DetachRolePolicy, events:RemoveTargets and friends only appear when CloudFormation has to clean up a half-built resource.

prowler and checkov agree: the deny-first policy scans clean; the admin-first policy is flagged for privilege escalation, resource exposure, and unrestricted infrastructure modification.

Learning from failure is strictly more informative than learning from success: a failure says what was needed, a success only says what happened.

Files

Output is split by lifetime. results/ holds the curated artifacts and is version-controlled; runs/ holds one directory per invocation and is not.

Path What it is
results/deploy-policy.json Record — 46 statements, one per action, Sids intact. Also the resume state
results/deploy-policy.compact.json Deployed — 15 statements, losslessly compacted
results/admin-policy.json Admin-first, all calls including reads
results/admin-policy.mutating.json Admin-first, reads dropped (comparable form)
results/admin-policy.converged*.json Admin-first after 10 deny-first passes
results/comparison.generated.md Measured tables — regenerated by compare_policies.py
results/comparison.md Hand-written analysis and verdict; never script-written
results/policy-diff.md Side-by-side, action by action
results/summary.md Action -> phase -> resources
results/admin-timing.json Wall-clock and action counts for the admin-first run
runs/<timestamp>/ Per-run logs (iterations.md, admin-run.md, admin-converge.md) and every streamed event (*-cloudtrail-raw.jsonl)
archive/ Artifacts from the earlier of the two accounts

Account IDs in every committed artifact are placeholders: 123456789012 for the account the current results came from and 210987654321 for the earlier one, kept distinct so the two-account history still reads correctly. The loop rewrites the account field of a seeded ARN to whatever is configured (debug category rehome), so results/deploy-policy.json still works as resume state despite naming an account that is not yours.

runs/ is gitignored: the logs are opened in append mode, so a tracked copy would be rewritten on every invocation. The per-run logs this write-up is based on are therefore not published; archive/ holds the curated remainder.

Prerequisites

cfn-lint and the aws CLI on PATH, credentials for the target account, and four IAM roles that the tools do not create. They only ever put-role-policy onto a role that already exists, so a missing one surfaces as a deploy failure rather than as a clear error.

All four are CloudFormation service roles: CloudFormation assumes them, so they trust cloudformation.amazonaws.com, not you.

cat > /tmp/cfn-trust.json <<'JSON'
{"Version": "2012-10-17", "Statement": [{
  "Effect": "Allow",
  "Principal": {"Service": "cloudformation.amazonaws.com"},
  "Action": "sts:AssumeRole"
}]}
JSON

# 1. deny-first subject: starts at ReadOnlyAccess, the loop grows it
aws iam create-role --role-name cfn-deploy-role \
    --assume-role-policy-document file:///tmp/cfn-trust.json
aws iam attach-role-policy --role-name cfn-deploy-role \
    --policy-arn arn:aws:iam::aws:policy/ReadOnlyAccess

# 2. admin-first subject: never fails, so CloudTrail records real calls
aws iam create-role --role-name cfn-admin-role \
    --assume-role-policy-document file:///tmp/cfn-trust.json
aws iam attach-role-policy --role-name cfn-admin-role \
    --policy-arn arn:aws:iam::aws:policy/AdministratorAccess

# 3. converge subject: bare on purpose, seeded from the admin-derived policy
aws iam create-role --role-name cfn-derived-role \
    --assume-role-policy-document file:///tmp/cfn-trust.json

# 4. cleanup: forces teardown of a stack the subject role cannot unwind
aws iam create-role --role-name cfn-cleanup-role \
    --assume-role-policy-document file:///tmp/cfn-trust.json
aws iam attach-role-policy --role-name cfn-cleanup-role \
    --policy-arn arn:aws:iam::aws:policy/AdministratorAccess

Role 3 is deliberately bare: converge_admin.py seeds it from results/admin-policy.json in order to measure that policy's shortfall, so any starting permissions would contaminate the measurement.

That trust policy carries no aws:SourceArn or aws:SourceAccount condition, which is the confused-deputy gap noted under Known limitations. Adding "Condition": {"StringEquals": {"aws:SourceAccount": "<account>"}} is strictly better and does not affect the experiment.

You need enough to create those roles, pass them to CloudFormation, and delete what a failed run abandons: iam:CreateRole, iam:PassRole, iam:PutRolePolicy, cloudformation:*, cloudtrail:LookupEvents, and delete permissions across the nine services purge_all() sweeps. In practice this is an admin in a sandbox account, which is the only place any of this should run — purge_all() deletes every resource whose name contains IAMD_STACK, and the deny-first loop deliberately drives the stack into failure states.

Running the tools

Copy .env.example to .env and set the two required values:

cp .env.example .env
$EDITOR .env          # IAMD_ACCOUNT and IAMD_STACK have no defaults

IAMD_STACK is deliberately not defaulted. It is the substring purge_all() matches orphaned resources on across nine services, so inheriting it would scope a real deletion to a name the operator never chose.

Configuration resolves from four sources, lowest precedence first: the defaults in tools/config.py, then .env, then the environment, then command-line flags. Every setting has all three spellings — IAMD_ADMIN_ROLE in the file or the environment is --admin-role on the command line — so a one-off run against a different target needs no file at all:

python3 tools/discover.py --account 123456789012 --region us-west-2 \
                          --stack my-stack --role my-deploy-role

--help on any tool lists the full set with its resolved defaults. --env-file points at a file somewhere other than the repo root, and pipeline.py exports whatever it resolved into the environment of every stage it runs, so a flag passed once applies to the whole workflow rather than to half of it.

Nothing that touches AWS starts without both: an ARN built against the wrong account is not an error the tools could recover from, since purge_all() deletes what it enumerates.

Reproducing the experiment

Use tools/pipeline.py rather than sequencing these by hand — it checks the preconditions each stage silently depends on and refuses instead of producing a plausible wrong answer:

python3 tools/pipeline.py --all --dry-run          # resolve the plan, run nothing
python3 tools/pipeline.py --stages preflight,admin,compare --allow-destructive
python3 tools/pipeline.py --all --allow-destructive

The three lines are increasing levels of commitment, and the middle one is the honest starting point: it exercises the admin-first half in about five minutes and writes a comparison, without spending hours on the deny-first loop.

Stage Wall clock What it costs
preflight seconds nothing; refuses rather than guesses
purge, reset ~3 min deletes leftovers from an abandoned run
deny hours (~20 deploy/fail iterations) the full deny-first derivation
admin ~5 min one deploy as admin, then CloudTrail
converge up to an hour measures the admin-first shortfall
compare seconds no AWS calls
teardown under a minute if clean, longer with a stack to delete deletes the stack and its orphans

Each stage rewrites only its own artifacts. Running the admin half alone regenerates results/admin-policy.* and leaves results/deploy-policy.* as it found them — the committed ones, in a fresh clone. The comparison then puts a fresh admin column beside a shipped deny-first column, which reads as a full reproduction and is not one. compare_policies.py states the age of each side at the top of its report so the two cannot be confused; deriving the deny-first column yourself means running the deny stage, and that is the one that costs hours.

Everything lands in runs/<timestamp>/ and overwrites results/, so git diff results/ after a run is the real comparison against what is committed here. A run is resumable: the loop seeds from results/deploy-policy.json, so an interrupted run costs only the current iteration. To measure a genuine from-zero derivation, include the reset stage — otherwise the seed makes it converge in one pass and measure nothing.

Expect the numbers to differ from the committed ones. AWS changes which denials surface and in what order, and the README records one such drift already: a from-zero rerun produced the same 44 actions in 17 statements rather than 15.

Every stage writes into one runs/<timestamp>/ and the outcome lands in pipeline.json, so a scheduler can branch on the result without parsing logs. Exit codes: 0 success, 1 a stage failed, 2 preflight refused, 3 bad invocation. Nothing that changes AWS state runs without --allow-destructive.

The individual tools remain runnable on their own:

python3 tools/discover.py          # deny-first: create -> rollback -> delete
python3 tools/admin_discover.py    # admin-first collection
python3 tools/compare_policies.py  # writes results/comparison.md
python3 tools/converge_admin.py    # measures the admin-first shortfall
python3 tools/uninstall.py         # teardown, harvesting delete permissions

Each invocation writes its logs and raw CloudTrail stream to a fresh runs/<timestamp>/, and overwrites the curated artifacts in results/. compare_policies.py reads from results/, so run it after both collectors.

Tests

Three tiers, each opt-in past the first, because a full deny-first run costs hours and real money and nothing that expensive belongs on the default path.

make test                      # unit: no AWS, no credentials, ~0.1s
make integration               # read-only against real AWS, seconds, free
make integration-destructive   # deploys and deletes one SNS topic, minutes

Unit covers the pure surface: precedence resolution and .env parsing, integer coercion, derived ARNs, rehome_account(), canon_resource(), the account guard, and — the one worth having — that compaction is lossless on every committed policy, which is the property the code itself falls back on and the README claims. Two parametrised checks also assert no real account ID has crept back into results/. This tier runs in CI.

Integration (read-only) proves the account guard against real STS: that it accepts matching credentials and refuses mismatched ones. That check is what stops purge_all() deleting in whichever account the credentials happen to name while every log line claims otherwise.

Integration (destructive) deploys tests/fixtures/selftest.yaml — a single SNS topic — through the real deploy(), asserts ProjectName reached CloudFormation and substituted into the resource name, then tears it down and confirms the stack is gone. It uses its own stack name (iam-discovery-selftest) on purpose: the purge and survey helpers match resources on a substring of the stack name, so a test sharing a name with the real stack could enumerate and delete it.

What no tier covers: the discovery loop itself. Whether harvest() extracts the right permission from a real denial is only answered by a full run against the real template, and that remains a manual exercise.

Log levels

IAMD_LOG_LEVEL sets console verbosity: ERROR, WARNING, INFO (default), DEBUG, TRACE. IAMD_DEBUG=1 is a shorthand for DEBUG.

Level What it adds
ERROR only failures that stop the run
WARNING lossy harvests, forced deletions, unscanned pushes
INFO phases, grants, purges, milestones — the default
DEBUG what each filter discarded, and why
TRACE every aws CLI invocation, its argv and exit code

INFO and above always reach runs/<timestamp>/iterations.md whatever the console threshold, because that file is the evidence trail — raising the threshold must never thin the record. DEBUG and TRACE go to runs/<timestamp>/debug.log so diagnostics never dilute it.

DEBUG covers thirteen decision points, not just discards:

Category The decision
errorcode an event whose error code is not treated as a denial
foreignrole an event belonging to another principal
noaction an event yielding no IAM action
unscoped a resource falling back to *
unparsed a failure reason matching no denial pattern
settle why CloudTrail polling stopped — early, or at the deadline
scanfinding Access Analyzer findings other than INVALID_ACTION
rewrite every silent ARN transformation canon_resource() performs
rehome a seeded ARN repointed from a placeholder account to yours
keptid a UUID-shaped ARN left scoped because its type is not per-deploy
widened specific ARNs discarded because * joined the set
notstuck a failing pass judged healthy, deferring recovery
readsplit an action classified as mutating by verb prefix

The discard trail is the one to reach for when a permission goes missing: every filter in the harvest path drops events silently, and a dropped event that mattered looks exactly like one that never did. Replaying this project's own recorded events showed events:DescribeRule reported as an opaque UnknownError 59 times and as a harvestable AccessDenied exactly once.

Runs are resumable: the loop seeds from results/deploy-policy.json, so a restart costs nothing but the current iteration. That path is deliberately not timestamped — moving it under runs/ would make every run start from scratch.

Edge cases handled, and why

Each of these was a silent failure first. In a loop that grants permissions until failures stop, anything that suppresses failure detection looks exactly like success — so every one of these produced a confident, wrong result before it was found.

Problem Handling
IAM/STS/global events invisible in-region Query us-east-2 and us-east-1
--max-results 50 silently truncating, oldest-first Full CLI auto-pagination
Stack events and CloudTrail each lossy, differently Union of both, tagged by source
...log-group:NAME:log-stream: never matches canon_resource() rewrites to :*
policy arn:... re-prefixed into a doubled ARN Strip everything before the second arn:
describe-stack-events returns whole history Time filter via calendar.timegm (DST-safe)
CloudTrail reports API names, not IAM actions Access Analyzer validate-policy gate
deploy returns before CloudFormation settles wait_terminal() on both sides
Failed pass with no new denials read as a stall Stall = no denials and no resource progress
DELETE_FAILED --retain-resources, then FORCE_DELETE_STACK
UPDATE_ROLLBACK_FAILED continue-update-rollback --resources-to-skip
Orphaned named resources -> AlreadyExists purge_all() across 9 services
Policy exceeding 10,240 bytes Lossless compaction, then managed-policy chunking
Generated probe template validate-template + cfn-lint before use

Compaction is lossless and proven so on every push. expand() flattens a policy to its (effect, action, resource) triples; the compact document must expand to the same set or the code falls back to the verbatim one. 46 statements -> 15, 12,577 -> 3,552 bytes, no grant widened. ARN wildcarding was rejected: it is lossy and larger (5,994 bytes).

Known limitations

  • Physical IDs cannot be scoped. lambda:*EventSourceMapping must use event-source-mapping:* — the UUID is minted per deploy, so a captured one matches exactly once. Scoping it caused a non-terminating loop.
  • The record and the deployed artifact differ. results/deploy-policy.json still contains s3:PutBucketEncryption and s3:PutBucketLifecycle because it records what was observed; only the compact artifact is scrubbed.
  • Account- and stack-specific. ARNs embed the account ID and stack name.
  • Trust policies were never examined. All three roles trust cloudformation.amazonaws.com with no aws:SourceArn/aws:SourceAccount condition. Prowler's iam_role_cross_service_confused_deputy_prevention would flag this.
  • iam:AttachRolePolicy does not constrain which policy is attached. Scoped to a role ARN, but the policy ARN is unconstrained. AWS's answer is a permissions boundary, which is not applied here.
  • The policy is not standalone. It contains zero read actions, because ReadOnlyAccess stays attached to the test role and masks every read denial. The deployed permission set is ReadOnlyAccess + these statements, and that managed policy grants read on essentially every service in the account. Deriving a self-contained policy means starting from no permissions at all.
  • Rollback discovered nothing. Two independent runs induced an update failure and harvested zero denials from it: every permission the rollback path needed had already been found during create. The three-phase structure earns its keep on delete, not on rollback.
  • Orphans in unpurged services stall the loop. AlreadyExists is not a denial, so it teaches the loop nothing while also blocking progress — the exact stall signature. Before events and stepfunctions were added to purge_all(), a stray EventBridge rule wedged a run at 21/23 and needed a manual delete-rule.
  • Rollback coverage is one induced failure, not an exhaustive exercise of every rollback path.

Not done

  • Auto-remap invalid actions using Access Analyzer's "Did you mean X" suggestion
  • Derive iam:PassRole from CloudTrail requestParameters.role
  • Auto-widen generated-ID ARNs
  • A hardening loop: scan -> remediate -> re-verify, with necessity testing by statement removal (would settle whether the kms:* grants are needed)
  • Permissions boundary and SCP guardrails, per AWS prescriptive guidance

License

Apache License 2.0 — see LICENSE and NOTICE.

The tools invoke the AWS CLI and cfn-lint as subprocesses rather than linking them, so no third-party license terms propagate to this work.

About

Deriving a least-privilege CloudFormation deployment policy: deny-first vs admin-first, measured. No LLM — every grant traceable to an observed denial.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages