Two experiments in working out exactly which IAM permissions the
order-pipeline stack needs in order to deploy, and a comparison of the two
methods.
Result: results/deploy-policy.compact.json — every grant traceable to an
observed denial and scoped to a specific ARN, with none falling back to *.
Verified to create, roll back, and delete the full 23-resource stack without
admin intervention.
The committed artifact is 15 statements / 3,552 bytes. A later from-zero rerun
produced the same 44 actions in 17 statements / 5,560 bytes: the loop captured
lambda:DeleteEventSourceMapping against a per-deploy UUID as well as the
wildcard, and the smaller figure reflects hand-scrubbing that is not in the
code. Treat 3,552 as a best case, not a reproducible output.
Deny-first (tools/discover.py). Start the CloudFormation service role at
ReadOnlyAccess, deploy, harvest every AccessDenied, grant exactly that action
on exactly that ARN, repeat until the stack completes. Then induce a rollback,
then delete. Create and delete each discover permissions the other does not;
rollback, in two separate runs, discovered none of its own — see Known
limitations.
Admin-first (tools/admin_discover.py). Give the role
AdministratorAccess, deploy once, and read CloudTrail for what it actually
called. This is what AWS productised as IAM Access Analyzer policy generation.
| Deny-first | Admin-first | |
|---|---|---|
| Cost | ~20 iterations, hours | 1 iteration, 292s |
| Mutating actions | 44 | 41 |
| Read actions | 0 (masked by ReadOnlyAccess) |
52 |
| Actions scoped to a real ARN | 44 | 17 |
Actions scoped to * |
0 | 24 |
| Deploys the stack | yes | no — 11 of 23 |
The action counts look comparable. They are not the interesting number.
Least privilege lives in the Resource field, and that is where the methods
diverge. A denial always names the resource that was refused. A successful
CloudTrail event usually has an empty resources array and no error text, so
there is nothing to scope to and the derivation falls back to *. The
admin-first policy is smaller because it is less specific, not tighter.
Making the admin-first policy work took 10 further deny-first passes plus two fixes it could not make itself — at which point it had become a deny-first run wearing admin-first's starting set.
Three things admin-first structurally cannot learn:
iam:PassRole. Authorised inside the calling API, so CloudTrail never logs it as its own event (documented). It halted the converge run at 21/23 until granted.- Correct action names. CloudTrail reports API names; for several S3
operations these differ from the authorising action
(
PutBucketEncryption->s3:PutEncryptionConfiguration). The static scan correctly rejects the API name, so the permission is lost rather than merely inert. - Failure and replacement paths.
lambda:RemovePermission,iam:DetachRolePolicy,events:RemoveTargetsand friends only appear when CloudFormation has to clean up a half-built resource.
prowler and checkov agree: the deny-first policy scans clean; the
admin-first policy is flagged for privilege escalation, resource exposure, and
unrestricted infrastructure modification.
Learning from failure is strictly more informative than learning from success: a failure says what was needed, a success only says what happened.
Output is split by lifetime. results/ holds the curated artifacts and is
version-controlled; runs/ holds one directory per invocation and is not.
| Path | What it is |
|---|---|
results/deploy-policy.json |
Record — 46 statements, one per action, Sids intact. Also the resume state |
results/deploy-policy.compact.json |
Deployed — 15 statements, losslessly compacted |
results/admin-policy.json |
Admin-first, all calls including reads |
results/admin-policy.mutating.json |
Admin-first, reads dropped (comparable form) |
results/admin-policy.converged*.json |
Admin-first after 10 deny-first passes |
results/comparison.generated.md |
Measured tables — regenerated by compare_policies.py |
results/comparison.md |
Hand-written analysis and verdict; never script-written |
results/policy-diff.md |
Side-by-side, action by action |
results/summary.md |
Action -> phase -> resources |
results/admin-timing.json |
Wall-clock and action counts for the admin-first run |
runs/<timestamp>/ |
Per-run logs (iterations.md, admin-run.md, admin-converge.md) and every streamed event (*-cloudtrail-raw.jsonl) |
archive/ |
Artifacts from the earlier of the two accounts |
Account IDs in every committed artifact are placeholders: 123456789012 for
the account the current results came from and 210987654321 for the earlier
one, kept distinct so the two-account history still reads correctly. The
loop rewrites the account field of a seeded ARN to whatever is configured
(debug category rehome), so results/deploy-policy.json still works as
resume state despite naming an account that is not yours.
runs/ is gitignored: the logs are opened in append mode, so a tracked copy
would be rewritten on every invocation. The per-run logs this write-up is
based on are therefore not published; archive/ holds the curated remainder.
cfn-lint and the aws CLI on PATH, credentials for the target account,
and four IAM roles that the tools do not create. They only ever
put-role-policy onto a role that already exists, so a missing one surfaces
as a deploy failure rather than as a clear error.
All four are CloudFormation service roles: CloudFormation assumes them, so
they trust cloudformation.amazonaws.com, not you.
cat > /tmp/cfn-trust.json <<'JSON'
{"Version": "2012-10-17", "Statement": [{
"Effect": "Allow",
"Principal": {"Service": "cloudformation.amazonaws.com"},
"Action": "sts:AssumeRole"
}]}
JSON
# 1. deny-first subject: starts at ReadOnlyAccess, the loop grows it
aws iam create-role --role-name cfn-deploy-role \
--assume-role-policy-document file:///tmp/cfn-trust.json
aws iam attach-role-policy --role-name cfn-deploy-role \
--policy-arn arn:aws:iam::aws:policy/ReadOnlyAccess
# 2. admin-first subject: never fails, so CloudTrail records real calls
aws iam create-role --role-name cfn-admin-role \
--assume-role-policy-document file:///tmp/cfn-trust.json
aws iam attach-role-policy --role-name cfn-admin-role \
--policy-arn arn:aws:iam::aws:policy/AdministratorAccess
# 3. converge subject: bare on purpose, seeded from the admin-derived policy
aws iam create-role --role-name cfn-derived-role \
--assume-role-policy-document file:///tmp/cfn-trust.json
# 4. cleanup: forces teardown of a stack the subject role cannot unwind
aws iam create-role --role-name cfn-cleanup-role \
--assume-role-policy-document file:///tmp/cfn-trust.json
aws iam attach-role-policy --role-name cfn-cleanup-role \
--policy-arn arn:aws:iam::aws:policy/AdministratorAccessRole 3 is deliberately bare: converge_admin.py seeds it from
results/admin-policy.json in order to measure that policy's shortfall, so
any starting permissions would contaminate the measurement.
That trust policy carries no aws:SourceArn or aws:SourceAccount
condition, which is the confused-deputy gap noted under Known limitations.
Adding "Condition": {"StringEquals": {"aws:SourceAccount": "<account>"}}
is strictly better and does not affect the experiment.
You need enough to create those roles, pass them to CloudFormation, and
delete what a failed run abandons: iam:CreateRole, iam:PassRole,
iam:PutRolePolicy, cloudformation:*, cloudtrail:LookupEvents, and
delete permissions across the nine services purge_all() sweeps. In
practice this is an admin in a sandbox account, which is the only place any
of this should run — purge_all() deletes every resource whose name
contains IAMD_STACK, and the deny-first loop deliberately drives the stack
into failure states.
Copy .env.example to .env and set the two required values:
cp .env.example .env
$EDITOR .env # IAMD_ACCOUNT and IAMD_STACK have no defaultsIAMD_STACK is deliberately not defaulted. It is the substring
purge_all() matches orphaned resources on across nine services, so
inheriting it would scope a real deletion to a name the operator never
chose.
Configuration resolves from four sources, lowest precedence first: the
defaults in tools/config.py, then .env, then the environment, then
command-line flags. Every setting has all three spellings — IAMD_ADMIN_ROLE
in the file or the environment is --admin-role on the command line — so a
one-off run against a different target needs no file at all:
python3 tools/discover.py --account 123456789012 --region us-west-2 \
--stack my-stack --role my-deploy-role--help on any tool lists the full set with its resolved defaults.
--env-file points at a file somewhere other than the repo root, and
pipeline.py exports whatever it resolved into the environment of every
stage it runs, so a flag passed once applies to the whole workflow rather
than to half of it.
Nothing that touches AWS starts without both: an ARN built against the
wrong account is not an error the tools could recover from, since
purge_all() deletes what it enumerates.
Use tools/pipeline.py rather than sequencing these by hand — it checks the
preconditions each stage silently depends on and refuses instead of producing
a plausible wrong answer:
python3 tools/pipeline.py --all --dry-run # resolve the plan, run nothing
python3 tools/pipeline.py --stages preflight,admin,compare --allow-destructive
python3 tools/pipeline.py --all --allow-destructiveThe three lines are increasing levels of commitment, and the middle one is the honest starting point: it exercises the admin-first half in about five minutes and writes a comparison, without spending hours on the deny-first loop.
| Stage | Wall clock | What it costs |
|---|---|---|
preflight |
seconds | nothing; refuses rather than guesses |
purge, reset |
~3 min | deletes leftovers from an abandoned run |
deny |
hours (~20 deploy/fail iterations) | the full deny-first derivation |
admin |
~5 min | one deploy as admin, then CloudTrail |
converge |
up to an hour | measures the admin-first shortfall |
compare |
seconds | no AWS calls |
teardown |
under a minute if clean, longer with a stack to delete | deletes the stack and its orphans |
Each stage rewrites only its own artifacts. Running the admin half alone
regenerates results/admin-policy.* and leaves results/deploy-policy.* as
it found them — the committed ones, in a fresh clone. The comparison then
puts a fresh admin column beside a shipped deny-first column, which reads as
a full reproduction and is not one. compare_policies.py states the age of
each side at the top of its report so the two cannot be confused; deriving
the deny-first column yourself means running the deny stage, and that is
the one that costs hours.
Everything lands in runs/<timestamp>/ and overwrites results/, so
git diff results/ after a run is the real comparison against what is
committed here. A run is resumable: the loop seeds from
results/deploy-policy.json, so an interrupted run costs only the current
iteration. To measure a genuine from-zero derivation, include the reset
stage — otherwise the seed makes it converge in one pass and measure nothing.
Expect the numbers to differ from the committed ones. AWS changes which denials surface and in what order, and the README records one such drift already: a from-zero rerun produced the same 44 actions in 17 statements rather than 15.
Every stage writes into one runs/<timestamp>/ and the outcome lands in
pipeline.json, so a scheduler can branch on the result without parsing logs.
Exit codes: 0 success, 1 a stage failed, 2 preflight refused, 3 bad
invocation. Nothing that changes AWS state runs without --allow-destructive.
The individual tools remain runnable on their own:
python3 tools/discover.py # deny-first: create -> rollback -> delete
python3 tools/admin_discover.py # admin-first collection
python3 tools/compare_policies.py # writes results/comparison.md
python3 tools/converge_admin.py # measures the admin-first shortfall
python3 tools/uninstall.py # teardown, harvesting delete permissionsEach invocation writes its logs and raw CloudTrail stream to a fresh
runs/<timestamp>/, and overwrites the curated artifacts in results/.
compare_policies.py reads from results/, so run it after both collectors.
Three tiers, each opt-in past the first, because a full deny-first run costs hours and real money and nothing that expensive belongs on the default path.
make test # unit: no AWS, no credentials, ~0.1s
make integration # read-only against real AWS, seconds, free
make integration-destructive # deploys and deletes one SNS topic, minutesUnit covers the pure surface: precedence resolution and .env parsing,
integer coercion, derived ARNs, rehome_account(), canon_resource(), the
account guard, and — the one worth having — that compaction is lossless on
every committed policy, which is the property the code itself falls back on
and the README claims. Two parametrised checks also assert no real account ID
has crept back into results/. This tier runs in CI.
Integration (read-only) proves the account guard against real STS: that
it accepts matching credentials and refuses mismatched ones. That check is
what stops purge_all() deleting in whichever account the credentials happen
to name while every log line claims otherwise.
Integration (destructive) deploys tests/fixtures/selftest.yaml — a
single SNS topic — through the real deploy(), asserts ProjectName reached
CloudFormation and substituted into the resource name, then tears it down and
confirms the stack is gone. It uses its own stack name (iam-discovery-selftest)
on purpose: the purge and survey helpers match resources on a substring of the
stack name, so a test sharing a name with the real stack could enumerate and
delete it.
What no tier covers: the discovery loop itself. Whether harvest() extracts
the right permission from a real denial is only answered by a full run against
the real template, and that remains a manual exercise.
IAMD_LOG_LEVEL sets console verbosity: ERROR, WARNING, INFO (default),
DEBUG, TRACE. IAMD_DEBUG=1 is a shorthand for DEBUG.
| Level | What it adds |
|---|---|
ERROR |
only failures that stop the run |
WARNING |
lossy harvests, forced deletions, unscanned pushes |
INFO |
phases, grants, purges, milestones — the default |
DEBUG |
what each filter discarded, and why |
TRACE |
every aws CLI invocation, its argv and exit code |
INFO and above always reach runs/<timestamp>/iterations.md whatever the
console threshold, because that file is the evidence trail — raising the
threshold must never thin the record. DEBUG and TRACE go to
runs/<timestamp>/debug.log so diagnostics never dilute it.
DEBUG covers thirteen decision points, not just discards:
| Category | The decision |
|---|---|
errorcode |
an event whose error code is not treated as a denial |
foreignrole |
an event belonging to another principal |
noaction |
an event yielding no IAM action |
unscoped |
a resource falling back to * |
unparsed |
a failure reason matching no denial pattern |
settle |
why CloudTrail polling stopped — early, or at the deadline |
scanfinding |
Access Analyzer findings other than INVALID_ACTION |
rewrite |
every silent ARN transformation canon_resource() performs |
rehome |
a seeded ARN repointed from a placeholder account to yours |
keptid |
a UUID-shaped ARN left scoped because its type is not per-deploy |
widened |
specific ARNs discarded because * joined the set |
notstuck |
a failing pass judged healthy, deferring recovery |
readsplit |
an action classified as mutating by verb prefix |
The discard trail is the one to reach for when a permission goes missing:
every filter in the harvest path drops events silently, and a dropped event
that mattered looks exactly like one that never did. Replaying this project's
own recorded events showed events:DescribeRule reported as an opaque
UnknownError 59 times and as a harvestable AccessDenied exactly once.
Runs are resumable: the loop seeds from results/deploy-policy.json, so a
restart costs nothing but the current iteration. That path is deliberately not
timestamped — moving it under runs/ would make every run start from scratch.
Each of these was a silent failure first. In a loop that grants permissions until failures stop, anything that suppresses failure detection looks exactly like success — so every one of these produced a confident, wrong result before it was found.
| Problem | Handling |
|---|---|
| IAM/STS/global events invisible in-region | Query us-east-2 and us-east-1 |
--max-results 50 silently truncating, oldest-first |
Full CLI auto-pagination |
| Stack events and CloudTrail each lossy, differently | Union of both, tagged by source |
...log-group:NAME:log-stream: never matches |
canon_resource() rewrites to :* |
policy arn:... re-prefixed into a doubled ARN |
Strip everything before the second arn: |
describe-stack-events returns whole history |
Time filter via calendar.timegm (DST-safe) |
| CloudTrail reports API names, not IAM actions | Access Analyzer validate-policy gate |
deploy returns before CloudFormation settles |
wait_terminal() on both sides |
| Failed pass with no new denials read as a stall | Stall = no denials and no resource progress |
DELETE_FAILED |
--retain-resources, then FORCE_DELETE_STACK |
UPDATE_ROLLBACK_FAILED |
continue-update-rollback --resources-to-skip |
Orphaned named resources -> AlreadyExists |
purge_all() across 9 services |
| Policy exceeding 10,240 bytes | Lossless compaction, then managed-policy chunking |
| Generated probe template | validate-template + cfn-lint before use |
Compaction is lossless and proven so on every push. expand() flattens a
policy to its (effect, action, resource) triples; the compact document must
expand to the same set or the code falls back to the verbatim one. 46 statements
-> 15, 12,577 -> 3,552 bytes, no grant widened. ARN wildcarding was rejected: it
is lossy and larger (5,994 bytes).
- Physical IDs cannot be scoped.
lambda:*EventSourceMappingmust useevent-source-mapping:*— the UUID is minted per deploy, so a captured one matches exactly once. Scoping it caused a non-terminating loop. - The record and the deployed artifact differ.
results/deploy-policy.jsonstill containss3:PutBucketEncryptionands3:PutBucketLifecyclebecause it records what was observed; only the compact artifact is scrubbed. - Account- and stack-specific. ARNs embed the account ID and stack name.
- Trust policies were never examined. All three roles trust
cloudformation.amazonaws.comwith noaws:SourceArn/aws:SourceAccountcondition. Prowler'siam_role_cross_service_confused_deputy_preventionwould flag this. iam:AttachRolePolicydoes not constrain which policy is attached. Scoped to a role ARN, but the policy ARN is unconstrained. AWS's answer is a permissions boundary, which is not applied here.- The policy is not standalone. It contains zero read actions, because
ReadOnlyAccessstays attached to the test role and masks every read denial. The deployed permission set isReadOnlyAccess+ these statements, and that managed policy grants read on essentially every service in the account. Deriving a self-contained policy means starting from no permissions at all. - Rollback discovered nothing. Two independent runs induced an update failure and harvested zero denials from it: every permission the rollback path needed had already been found during create. The three-phase structure earns its keep on delete, not on rollback.
- Orphans in unpurged services stall the loop.
AlreadyExistsis not a denial, so it teaches the loop nothing while also blocking progress — the exact stall signature. Beforeeventsandstepfunctionswere added topurge_all(), a stray EventBridge rule wedged a run at 21/23 and needed a manualdelete-rule. - Rollback coverage is one induced failure, not an exhaustive exercise of every rollback path.
- Auto-remap invalid actions using Access Analyzer's "Did you mean X" suggestion
- Derive
iam:PassRolefrom CloudTrailrequestParameters.role - Auto-widen generated-ID ARNs
- A hardening loop: scan -> remediate -> re-verify, with necessity testing by
statement removal (would settle whether the
kms:*grants are needed) - Permissions boundary and SCP guardrails, per AWS prescriptive guidance
Apache License 2.0 — see LICENSE and NOTICE.
The tools invoke the AWS CLI and cfn-lint as subprocesses rather than
linking them, so no third-party license terms propagate to this work.