AWS FinOps AI agents IAM

AWS FinOps agent: AccessDenied when the agent tries to act on a cost finding

Fix AccessDenied when an AWS FinOps agent tries to act on a cost finding (stop an instance, delete a volume, buy a commitment) by separating read-first observation from a narrowly scoped, human-gated write path.

Jerzy Kopaczewski ·
Your FinOps agent reads cost and usage data perfectly - it summarises the bill, flags idle resources, spots anomalies - but the moment it tries to act on a finding (stop an instance, delete an unattached volume, buy a Savings Plan) it fails with AccessDenied. In most cases this is not a bug to fix by widening permissions. It is the read-first design working as intended, and the right response is a deliberately scoped, human-gated write path - not a broad grant.

This runbook covers IAM access failures when a FinOps agent moves from observing to acting. For how the agent fits together (observe / recommend / act) and why the runbooks behind it matter, see AWS FinOps agent: cost, function, and runbooks. To design the scoped-write and approval layer safely, book a consulting session.

Symptoms

An error occurred (AccessDeniedException) when calling the <Operation>
operation: User: arn:aws:sts::...:assumed-role/finops-agent-role is not
authorized to perform: <action> (e.g. ec2:StopInstances,
ec2:DeleteVolume, savingsplans:CreateSavingsPlan) on resource: <arn>
because no identity-based policy allows the <action> action

Observable impact:

  • The agent’s read and analysis steps all succeed - Cost Explorer, CUR queries, recommendations render fine
  • Only the action step fails, and only for write/mutate operations
  • The failure is consistent (every write denied), not intermittent - a signal it is a policy boundary, not a transient error

Cause

A well-built FinOps agent is given read-first IAM: broad read access to cost and usage data, and little or no write access by default. That is the correct posture. An agent that can see the bill is useful; an agent that can silently change infrastructure to lower it is a production risk. So an AccessDenied on an action is usually the design holding the line, not a misconfiguration.

The failure resolves into one of three situations, and only one of them is a plain permission gap:

  1. Read-first by design (the common case). The execution role intentionally omits write actions. The agent should be recommending this change for a human to approve, not performing it. The fix is a gated write path, not a wider role.
  2. A genuinely intended action is missing its permission. You have decided this specific, bounded action should be automated (for example, stopping instances tagged auto-stop: true on a schedule). The single action is simply not yet on the role.
  3. Identity or boundary mismatch. A permissions boundary, SCP, or the wrong assumed role is blocking the action even though the identity-based policy looks correct.

Common triggers:

  • Write action never granted - the role has ce:*, cur:*, cloudwatch:Get* read scope but no ec2:StopInstances / ec2:DeleteVolume / savingsplans:CreateSavingsPlan
  • Permissions boundary caps the role - the action is in the policy but denied by a boundary or an Organizations SCP
  • Wrong role in effect - the agent assumed its read role for a step that was meant to assume a separate, narrowly scoped write role
  • Resource/condition too tight - the action is allowed but the resource ARN or a condition key (tag, region) does not match the target

Fix

Step 1: Decide whether this action should be automated at all

Before touching IAM, answer the governance question: should the agent perform this action autonomously, or recommend it for a human? Deleting volumes, stopping instances, and buying multi-year commitments carry real blast radius. The default for anything consequential is recommend, with a human gate - which means the AccessDenied is correct and the fix lives in your approval workflow, not the role.

Only proceed to Step 2 if you have deliberately decided this specific, bounded action is safe to automate.

Step 2: Grant the single action, scoped to the exact resource

If the action is genuinely intended, add just that action - never a wildcard - and constrain it by resource ARN and, where possible, a tag or region condition so the agent can only touch resources explicitly opted in.

{
  "Version": "2012-10-17",
  "Statement": [{
    "Sid": "FinOpsAgentScopedStop",
    "Effect": "Allow",
    "Action": "ec2:StopInstances",
    "Resource": "arn:aws:ec2:*:111122223333:instance/*",
    "Condition": {
      "StringEquals": { "aws:ResourceTag/finops-agent-managed": "true" }
    }
  }]
}

Prefer a separate narrow write role that the agent assumes only for the approved action, keeping the default execution role read-only. Do not add write scope to the broad read role.

Step 3: Check for a boundary or SCP if the grant does not take effect

If the identity-based policy now allows the action but it still denies, a permissions boundary or an Organizations SCP is capping the role. Confirm the effective permissions rather than reading the attached policy alone.

# Simulate the exact action against the exact resource, as the agent's role
aws iam simulate-principal-policy \
  --policy-source-arn arn:aws:iam::111122223333:role/finops-agent-write-role \
  --action-names ec2:StopInstances \
  --resource-arns arn:aws:ec2:eu-west-1:111122223333:instance/i-0abc123 \
  --region us-east-1

If the simulation shows implicitDeny with a boundary or SCP named, the fix is at that layer, not the role policy.

Step 4: Wire the action behind a human gate

For any action with production blast radius, the grant should sit behind an approval step so the agent proposes and a human confirms. This is the pattern from the FinOps agent article: the agent observes and analyses, the runbook encodes the safety checks, and a human approves anything that mutates infrastructure. An agent wired straight to write permissions skips exactly the step that keeps you out of an incident.

Validation

Confirm the agent can perform the one approved action on an opted-in resource, and still cannot touch anything else.

# Should now succeed for an explicitly tagged, opted-in resource
aws sts get-caller-identity --region us-east-1
# Confirm the write role is in effect, then exercise the approved action
# against a test resource carrying finops-agent-managed=true and verify success.

# Negative check: the same action on a NON-tagged resource must still be denied
aws iam simulate-principal-policy \
  --policy-source-arn arn:aws:iam::111122223333:role/finops-agent-write-role \
  --action-names ec2:StopInstances \
  --resource-arns arn:aws:ec2:eu-west-1:111122223333:instance/i-0notmanaged \
  --region us-east-1
# Expect: implicitDeny - the scope holds, the agent only touches opted-in resources.

Expected end state: the agent reads cost data as before, performs the single approved action on opted-in resources through the gated path, and is still denied every other write - the read-first posture intact.