AWS FinOps agent: what it does, what it costs, and the runbooks behind it

Jerzy Kopaczewski 20 September 2026 12 min read
Contents

AWS FinOps agent: what it does, what it costs, and the runbooks behind it

A FinOps agent is an AI agent pointed at your AWS bill: it reads cost and usage data, spots waste and anomalies, and proposes action. The appeal is obvious when cloud spend is sprawling and nobody has time to chase every line item. The catch is equally obvious to anyone who runs production - an agent that can see your bill is useful; an agent that can change your infrastructure to cut it is a risk. This article explains what an AWS FinOps agent actually does, what it costs to run, and why the runbooks behind it are what separate a safe one from an expensive mistake.

“FinOps agent” is one of a cluster of operational AI-agent ideas gaining traction on AWS, alongside DevOps agents and security agents. They share a shape: an agent with read access to a domain (cost, pipelines, security posture), a reasoning loop, and some ability to act. FinOps is the one with the clearest commercial case, because the thing it reasons about - spend - is measured in money, every day. So it is the one worth getting right first.

The framing here is the same production-first lens we bring to the rest of our agentic-AI writing: an agent is only as safe as the guardrails around what it can do. A FinOps agent that recommends is a reporting tool with a brain. A FinOps agent that acts is a change-management system, and it needs to be treated like one.

 

Thinking about pointing an AI agent at your AWS bill?

Book a free 30-min call

 

What an AWS FinOps agent actually does

Strip away the branding and a FinOps agent does three things, in increasing order of risk:

CapabilityWhat it meansRisk level
ObserveReads Cost Explorer, CUR, CloudWatch, and tagging data; summarises spend, trends, and anomalies in plain languageLow - read-only
RecommendProposes concrete actions: rightsizing, idle-resource cleanup, commitment purchases, storage-class changesLow to medium - advice, not action
ActExecutes changes: stops instances, resizes, buys Savings Plans, deletes unattached volumesHigh - production blast radius

Most of the value is in the first two tiers. An agent that reliably answers “why did this month’s bill jump” and “what are my ten biggest savings opportunities right now” already pays for itself, because that analysis is tedious and usually under-resourced. The third tier - acting automatically - is where teams get into trouble, and it is entirely optional. The design decision that matters most is not “how smart is the agent” but “what is it allowed to do without a human.”

 

Where it fits in agentic AI on AWS

A FinOps agent is not a standalone product you buy so much as a pattern you build, usually on the same agent infrastructure as everything else. On AWS that typically means Amazon Bedrock AgentCore for the runtime, memory, and gateway, with the agent given scoped, read-first access to your cost and usage data. If you are weighing whether to build on AgentCore, self-host, or use a framework, we compared those paths in AgentCore vs self-hosted vs frameworks.

It sits alongside two sibling agents in the same family:

  • DevOps agent - reasons about pipelines, deployments, and infrastructure drift.
  • Security agent - reasons about posture, findings, and misconfigurations.

The three overlap in how they are built and governed, which is why it is worth getting the first one right as a template. FinOps leads because its output is denominated in money, so the return is easy to measure.

 

What a FinOps agent costs to run

The irony is not lost on anyone: an agent whose job is to cut cost has a cost of its own. It is small relative to a sprawling bill, but you should model it rather than assume it away. The components:

  • Model inference - every analysis is one or more LLM calls. A daily cost summary across a large account is cheap; an agent re-reasoning over granular CUR data on every query is not. The usage pattern drives this far more than the per-token rate.
  • Agent runtime - if you build on AgentCore, there is a runtime and gateway cost, plus the OpenSearch Serverless floor for memory (on the order of a few hundred dollars a month before you process a single query). We broke that floor down in the AgentCore pricing piece.
  • Data access - Cost Explorer API calls and CUR queries (via Athena) carry their own small charges that add up at high query frequency.

The practical rule: a FinOps agent is cheap when it runs scheduled, bounded analyses (a daily digest, a weekly opportunity report) and expensive when it is wired to re-reason continuously. Design the cadence deliberately. The agent should save multiples of what it costs, or it is just a more interesting way to spend money.

 

The runbooks are the actual product

Here is the part most “AI will optimise your cloud bill” pitches skip. A recommendation is not a result. “Delete these 40 unattached EBS volumes” is only safe if someone knows those volumes are genuinely orphaned and not a detached-but-needed snapshot source. The gap between a finding and a safe action is a runbook, and that is where the real engineering lives.

A FinOps agent earns trust by being wired to runbooks rather than to raw API permissions:

  • Rightsizing - the agent flags an over-provisioned instance; the runbook defines how to verify utilisation over a representative window, check for scheduled peaks, and resize with a rollback path. Not every low-CPU instance is waste.
  • Idle-resource cleanup - the agent lists unattached volumes, idle load balancers, orphaned snapshots; the runbook defines the “is this really safe to delete” checks and a quarantine step before deletion.
  • Commitment purchases - the agent models Savings Plans or Reserved Instance coverage; the runbook keeps the actual purchase a human decision, because it is a one-to-three-year financial commitment. We cover that trade-off in Savings Plans vs Reserved Instances.
  • Anomaly response - the agent detects a spend spike; the runbook distinguishes a real anomaly from a billing artefact (a one-off data-transfer charge, a reserved-instance amortisation) before anyone is paged.

The pattern across all four: the agent does the observation and the analysis, the runbook encodes the judgement, and a human approves anything with real blast radius. An agent wired straight to write permissions skips the middle step, which is exactly the step that keeps you out of an incident.

 

Keeping the agent itself governed

One FinOps agent is a tool. A FinOps agent, plus a DevOps agent, plus a security agent, plus whatever three teams built independently, is the cloud sprawl story repeating itself one layer up - agents multiplying with inconsistent permissions and no shared oversight. The governance that keeps a FinOps agent safe is the same governance that keeps the whole agent estate safe:

  • Scoped, read-first IAM - the agent reads cost and usage data by default; write permissions are the exception, narrowly scoped, and audited.
  • A human gate on anything that acts - approval before any change with production blast radius.
  • Provenance and audit - every recommendation and action logged, so a change can be traced back to the agent and the data that prompted it.
  • An owner - one team accountable for the agent’s scope and behaviour, not a tool that quietly accreted permissions.

This is the same discipline we apply to showback and cost accountability generally. A FinOps agent is a faster way to produce the insight; it does not replace the operating model around it, which we describe in implementing a showback model.

 

How we can help

At Devopsity, FinOps and agentic AI are both core to what we do, which makes the FinOps agent a natural fit. We help teams stand one up the safe way: scoped read-first access to cost and usage data, the analysis that surfaces real savings, and - the part that matters - the runbooks that turn a finding into a safe action with a human gate where the blast radius is real. We build it on the right foundation (AgentCore or an alternative, depending on your estate) and keep the agent itself governed so it does not become another sprawl problem.

If your AWS bill is growing faster than anyone can explain and you are wondering whether an agent is the answer, let’s talk about your cloud costs.

Jerzy Kopaczewski

Considering a FinOps agent for your AWS account?

Book a free 30-minute call. No pitch - a technical conversation about where an agent helps on cost, what it would cost to run, and the runbooks and guardrails that keep it safe to let near production.

Book a call

Frequently asked questions

What is an AWS FinOps agent?

It is an AI agent with access to your AWS cost and usage data (Cost Explorer, CUR, CloudWatch, tagging) that summarises spend, detects anomalies, and recommends savings. More capable versions can also act on infrastructure, but acting automatically is optional and carries production risk - the safest versions observe and recommend, with a human approving any change.

How much does a FinOps agent cost to run?

The main components are model inference (LLM calls per analysis), agent runtime (if built on Bedrock AgentCore, including the OpenSearch Serverless floor of a few hundred dollars a month), and data-access charges (Cost Explorer and Athena queries). It is cheap when it runs bounded, scheduled analyses and expensive when wired to re-reason continuously. It should save multiples of what it costs.

Should a FinOps agent make changes automatically?

Usually not without a human gate. Observing and recommending is low-risk and captures most of the value. Acting automatically - stopping instances, deleting volumes, buying commitments - has real production and financial blast radius, so the safe pattern is to wire the agent to runbooks with human approval for anything consequential.

Why are runbooks so important for a FinOps agent?

Because a recommendation is not a safe action. The gap between “this looks like waste” and “it is safe to delete this” is judgement, and a runbook encodes that judgement: how to verify utilisation, how to confirm a resource is truly orphaned, how to distinguish a real anomaly from a billing artefact. Without runbooks, an agent with write access is an incident waiting to happen.

Do I need Bedrock AgentCore to build one?

No, but it is the common AWS-native path for the runtime, memory, and gateway. You can also self-host or use a framework. The right choice depends on your estate; the agent’s safety comes from scoped IAM and runbooks, not from the specific runtime.

AWS FinOps AI agents agentic AI cost optimisation Bedrock AgentCore cloud cost management

Read also:

Previous post Next post