Guide to Kubernetes Cost Management - Part 3
Rajith Attapattu
September 5, 2026 · 12 min read

In short
Once a team can find optimization opportunities, human attention becomes the bottleneck. Applying the FinOps Crawl, Walk, Run model to Kubernetes, this piece argues that maturity means shortening the distance between detecting an inefficiency and safely acting on it, with SRE Agents doing the investigation inside boundaries humans set.
From Optimization to a Repeatable Practice
Part 1 of this series asked where Kubernetes spend actually comes from. Part 2 asked how you find and validate optimization opportunities once you can see that spend clearly.
This part asks a different question: how do you turn that into something an organization can actually sustain?
Finding one rightsizing opportunity isn’t particularly hard. Give an engineer a dashboard, a few weeks of utilization data and some time, and they will find a workload that is oversized. The harder problem is doing that continuously, across hundreds or thousands of workloads, several clusters and multiple engineering teams, without it turning into a full-time job for someone.
That’s an organizational problem as much as a technical one. It raises questions that don’t have purely technical answers:
- Who owns cost for a given workload?
- Who investigates an anomaly when one shows up?
- Who decides whether a recommendation is safe to act on?
- Who actually makes the change?
- Who checks whether it worked?
In practice, most organizations don’t answer these questions all at once. They move through stages, and the FinOps community already has a reasonably good name for that progression: Crawl, Walk, Run.
FinOps Maturity Is More Than Better Dashboards
I don’t want to spend too long on the generic version of the FinOps maturity model, it’s well documented elsewhere. What I think is more useful is what it actually looks like when applied to Kubernetes specifically, because the technical maturity and the organizational maturity have to develop together.
Roughly:
Crawl, the primary question is where is our money going. Capabilities are mostly about visibility: cluster and namespace cost, ownership, historical trends, obvious waste. Humans do almost all of the investigation.
Walk, the question shifts to where can we improve. This is where rightsizing recommendations, dormant workload detection, cluster efficiency, cost anomalies and showback start to appear. Systems increasingly surface opportunities, but humans still evaluate and execute most of the changes.
Run, the question becomes how do we keep doing this continuously, within guardrails we’ve already agreed on. This is where automated investigation, policy-aware recommendations, approved runbooks and post-change validation start to matter, and where some actions can be executed without a human triggering each one individually.
That progression leads to what I think is the central idea of this article:
Maturity isn't simply having more cost data. It is reducing the distance between detecting an inefficiency and safely doing something about it.
A lot of organizations get stuck at Crawl not because their dashboards are bad, but because closing that distance requires trust, and trust takes longer to build than a dashboard does.
Crawl: Make the Data Trustworthy First
Before an organization can move faster, the underlying data needs to hold up under scrutiny. This is where I think a lot of cost initiatives quietly stall, not because nobody built the dashboard, but because engineers don’t trust the numbers on it.
A few things need to be true first.
Allocation. Can you reliably attribute infrastructure cost through Cluster → Namespace → Workload → Pod/Container? If two people run the same query and get different numbers, nobody is going to act on either of them.
Ownership. Can that same cost be mapped through Application → Team → Business Unit? Attribution without ownership just produces a report nobody feels responsible for.
Historical context. Can you tell whether today’s spend and utilization are normal for this workload, or is every number evaluated in isolation?
Cloud and Kubernetes context. Can what you’re being billed for by the cloud provider actually be connected to the workloads consuming it, rather than living in two separate systems that never quite agree?
None of this is glamorous, and none of it produces a recommendation an engineer can act on today. But at this stage, that isn’t the objective. The objective is getting engineering and FinOps to the point where they can look at the same number and have the same conversation about it. Skip this step and every later stage inherits the distrust.
Walk: Move From Reporting to Recommendations
Once the data holds up, the next step is moving from describing cost to identifying what should actually happen next.
There’s a real difference between:
“Payments costs $18,000 per month.”
and:
“These three workloads account for most of that cost, two appear significantly oversized relative to what they use, and one hasn’t shown meaningful traffic in 30 days.”
The first is a fact. The second is something an engineer can act on. Getting from one to the other is most of what Walk is about, and it shows up in a few familiar shapes:
Workload rightsizing. Historical CPU and memory behavior turns into a recommendation with an estimated saving, the approach we covered in detail in Part 2.
Dormant workloads. Something that looks unused gets flagged, investigated, and either remediated or confirmed as intentionally idle.
Cluster and node efficiency. Workload requirements, scheduling behavior and node utilization point toward a better-fitting infrastructure configuration.
Cost anomalies. A baseline exists, something deviates from it, and that deviation gets investigated rather than discovered a month later on an invoice.
There’s an important distinction here that we made in Part 2 and that becomes even more relevant in this article: a recommendation without operational context is incomplete. Knowing a workload is oversized isn’t the same as knowing it’s safe to resize. That requires utilization data, but also application performance, error rates, incident history and recent deployments. This is the point where cost optimization and observability stop being separate disciplines and start being the same investigation.
The Recommendation Bottleneck
Here’s where I think the article needs to take a turn, because this is where a lot of FinOps tooling quietly runs out of road.
Suppose your cost platform is mature enough to surface 300 potential rightsizing opportunities across your clusters this month. That sounds like a win. It also creates a new problem, because for each of those 300, someone now has to:
- decide which of them are actually credible
- understand how that specific application behaves
- judge whether the change is safe
- make the configuration change
- deploy it
- watch what happens
- confirm the saving actually materialized
Do that seven times and it’s a good afternoon. Do it 300 times a month, every month, across a growing set of clusters, and it stops being realistic for any team to keep up with by hand.
Human attention eventually becomes the bottleneck. Not the ability to generate recommendations, the ability to act on them responsibly. That’s the problem the next section is actually about.
From Recommendations to Agent-Assisted FinOps
I want to be careful about how I introduce this, because it’s easy to overstate. The right starting question isn’t “can AI optimize your Kubernetes infrastructure.” It’s narrower and more useful than that: what if some of the investigation an SRE would normally do by hand could happen automatically, before it ever reaches a human?
An SRE Agent’s value isn’t that it can generate text about a recommendation. It’s that it can reason across several categories of context that usually live in separate systems and get correlated manually, if they get correlated at all.
Cost context: historical spend, allocation, estimated savings, anomalies.
Kubernetes context: requests and limits, replica counts, workload configuration, node utilization, scheduling behavior, deployment history.
Observability context: CPU and memory behavior, latency, errors, availability, application performance.
Engineering context, where it’s available: recent deployments, relevant Git changes, approved runbooks, organizational policy.
Pulling one of these signals is useful. Correlating all of them against a specific workload before a human ever looks at it is where this actually starts to save time rather than just generating another dashboard tile. Let’s make this concrete with an example, because the argument above is easy to agree with in the abstract and much more interesting when you follow one recommendation all the way through.
Following One Recommendation Through the Loop
Say we have a service called checkout-api.
Detect. The system notices each pod requests 2 CPU, but observed p95 CPU usage over a reasonable evaluation window sits around 700m. That’s a meaningful gap, and it’s been stable, not a one-off dip. There’s an estimated monthly saving attached to it.
Investigate. This is where it gets more interesting than a static rightsizing report. Instead of stopping at “this looks oversized,” the agent pulls historical CPU and memory, replica history, traffic patterns, latency, error rate, recent incidents and recent deployments for checkout-api specifically. It’s trying to answer one question: is this apparent opportunity consistent with how the application actually behaves, or is there a reason the request is set where it is?
Recommend. Say everything checks out. The recommendation isn’t just “lower the request.” It’s something closer to: reduce the CPU request from 2 CPU to 1.25 CPU, while keeping enough headroom above observed peaks to absorb normal variance. That number should come from the workload’s actual behavior and the organization’s own risk tolerance, not from an LLM picking a number that looks reasonable.
Human guidance. This is where organizational policy enters the picture. Maybe the policy is that all production rightsizing changes require explicit approval. Maybe it’s that changes meeting a defined, low-risk profile can proceed through an already-approved workflow. Either way, a human or a human-defined policy decides what happens next, not the agent.
Act. Depending on that policy, the path might be: generate the configuration change, open a pull request, wait for approval, deploy through the existing GitOps pipeline. Or, for a workload that already qualifies under an approved runbook, execute that runbook directly.
Validate. The change being deployed isn’t the finish line. After it ships, the same signals get watched again: CPU utilization, throttling, memory, latency, error rate, availability, and yes, cost. A good outcome looks like: resource allocation dropped, application performance stayed within expected thresholds, and infrastructure cost went down. A bad outcome looks like: latency moved outside the accepted range, recommend reverting.
That’s the loop, end to end: Detect → Investigate → Recommend → Approve → Act → Validate. The recommendation was the easy part. Everything after it is where the actual engineering judgment lives, and it’s also exactly the part that’s hardest to scale by hand.
Human SREs Define the Boundaries
I don’t think the goal here is removing the SRE from the loop. I think the more accurate way to describe it is: move the SRE up the decision hierarchy.
Instead of manually investigating every candidate opportunity, human SREs and platform teams spend their time defining the boundaries the agent operates within: what it’s allowed to look at, which tools it can use, which actions it can take, which runbooks are pre-approved, which environments it’s allowed to touch, what counts as an acceptable risk, when approval is mandatory, how success is validated, and under what conditions a change gets rolled back automatically.
In practice, that tends to show up as a few distinct models rather than one:
Recommendation only. Agent investigates and recommends, a human does everything else.
Human-in-the-loop. Agent investigates, proposes a specific action, a human approves it, then the agent (or the existing pipeline) executes and validates.
Policy-controlled automation. Agent investigates, executes through an already-approved workflow, validates the result, and escalates to a human only when something falls outside the policy.
None of these is inherently more mature than the others. The objective isn’t maximum autonomy. It’s the appropriate level of autonomy for the risk involved, and that’s going to differ by workload, by environment and by organization.
Run: Closing the Loop
Coming back to the maturity model, this is what Run actually looks like end to end, not as a single automated action but as a continuous cycle:
Observe cost and operational behavior continuously, not once a month. Detect anomalies, dormant resources and optimization opportunities as they appear. Investigate by pulling in cost, Kubernetes and observability context together. Recommend a specific action grounded in that context. Govern by applying whatever organizational policy and approval that action requires. Act by executing through an approved path. Validate both the cost outcome and the application’s behavior afterward. Learn by feeding that outcome back into how future recommendations get made.
Observe → Detect → Investigate → Recommend → Govern → Act → Validate → Learn.
That’s a longer loop than the one we ended Part 2 with, and deliberately so. Govern and Learn are the two additions, and they’re the two steps that separate a team occasionally automating a fix from an organization running a repeatable practice.
FinOps and SRE Are Converging
There’s an organizational shift underneath all of this that I think is worth naming directly.
Traditionally, these have been three separate conversations. FinOps asks what we’re spending. Engineering asks how the application is performing. SRE asks whether we can operate it reliably. For a lot of infrastructure decisions, that separation worked fine.
It stops working cleanly for Kubernetes, because the same decision usually touches all three at once. Reducing a CPU request affects cost and performance together. Reducing replica count affects cost and availability together. Changing node types affects cost, scheduling and application behavior together. Even removing a dormant workload requires understanding whether it actually serves an operational purpose before you can call it waste.
Cost becomes another operational signal, alongside latency, error rate and availability, not a separate report that shows up at the end of the month. The mature version of this isn’t FinOps handing engineering a spreadsheet of savings opportunities. It’s financial and operational context sitting next to each other when the decision is actually being made.
Maximum Automation Isn’t the Goal
I want to push back on one assumption before closing, because it’s an easy one to make once you’ve read the last few sections: that a mature FinOps practice is one where everything eventually gets automated.
It isn’t, and I don’t think it should be. A mature financial institution might reasonably require human approval on every production change, full stop. A mature SaaS company might automate a class of low-risk changes without a second thought. The same organization might apply completely different policies depending on environment and workload, and all of the following can coexist inside one mature practice:
A dormant workload in development gets shut down automatically. A rightsizing recommendation in a non-production namespace gets applied through a policy-controlled workflow without a human in the loop. The same recommendation against a production stateless service turns into a pull request that waits for a reviewer. A recommendation against a production database turns into investigation and a report, and nothing more, because the blast radius of getting it wrong is too high to automate away.
FinOps maturity isn’t about removing humans. It’s about knowing which decisions require humans and which can safely be automated, and having the judgment, and the policy, to tell the two apart consistently.
Understand, Optimize, Operationalize
Pulling the three parts of this series together:
Part 1 was about understanding where Kubernetes cost actually comes from, and why that turns out to be harder than reading a cloud bill.
Part 2 was about turning that visibility into a continuous feedback loop: observing behavior, identifying inefficiency, making recommendations grounded in real usage, validating the outcome.
Part 3 has been about what it takes to run that loop as an organizational practice rather than a project someone picks up every few months: trustworthy data first, recommendations that carry real context, and increasing amounts of investigation and execution handled by an agent operating inside boundaries that humans set deliberately.
That last part matters more than the automation itself. The end state I’d argue for isn’t an agent that autonomously trims infrastructure until the bill looks better. It’s an organization where cost, utilization, performance and reliability are understood together continuously, where humans set the objectives, the policies and the boundaries, and where automation absorbs an increasing share of the repetitive investigation and execution that used to consume an engineer’s afternoon.
That’s where I think Kubernetes FinOps is heading.
Related reading
- Guide to Kubernetes Cost Management - Part 1: where Kubernetes cost actually comes from.
- Guide to Kubernetes Cost Management - Part 2: turning visibility into a continuous optimization loop.
- SRE Agent (Raiya): the agent-assisted investigation and remediation workflow this article builds toward.
- Cost & FinOps: attribution, rightsizing, and chargeback without assembling this yourself.
