Cloud Intelligence™
AI Cost Management: How Gartner's CFM Evaluation Criteria Are Catching Up to AI Spend
AI cost management applies Gartner's CFM criteria, financial risk, forecasting, efficiency, accountability, to token and GPU spend.
This page is also available in Deutsch, Español, Français, Italiano, 日本語, and Português.
About Josh Palmer
I'm Josh Palmer, Head of Content at DoiT, where I split my time across multiple business units including DoiT Cloud Intelligence, PerfectScale (Kubernetes cost optimization), and SELECT (Snowflake, Databricks, and BigQuery cost optimization). Before DoiT, I spent four and a half years at OnBoard building content for a board intelligence platform used by 6,000+ organizations, and before that, two years as Content Marketing Manager at Zylo, a SaaS management platform.
My personal pageTL;DR: AI cost management is the practice of measuring, attributing, forecasting, and optimizing what an organization spends on AI and LLM workloads: tokens, GPU compute, inference, and the infrastructure around them. It's not a separate discipline from cloud financial management (CFM) so much as an extension of it.
What is AI cost management?
AI cost management is the set of practices and tools that connect what an organization spends on AI, model API calls, GPU infrastructure, agentic workloads, back to the teams, products, and outcomes driving that spend. The goal is the same one FinOps set for cloud infrastructure a decade ago: give finance and engineering a shared, accurate picture of spend so they can make decisions from the same numbers, instead of finance seeing a bill and engineering seeing a black box.
The reason AI cost management gets discussed as its own category, rather than just "cloud FinOps, but for AI," comes down to how differently AI infrastructure behaves. Cloud spend is relatively stable: an instance has an owner, a bill has a resource ID, a tag survives the trip from provisioning to invoice. AI workloads break several of those assumptions at once. A shared model API account can serve a dozen teams from a single billing line. An LLM gateway can strip caller identity before a request ever reaches the provider. An agentic pipeline can spawn sub-agents overnight that trigger real infrastructure costs no one instrumented for, because no one anticipated the call pattern.
That's the gap AI cost management exists to close: applying the discipline of cost allocation, forecasting, optimization, and governance to workloads that move faster and share more than the infrastructure FinOps was originally built for.
Why is Gartner's cloud financial management criteria expanding to cover AI?
Gartner's Magic Quadrant for Cloud Financial Management Tools already asks whether a vendor can do four things: manage financial risk, forecast spend, increase efficiency, and increase accountability. None of those mandatory capabilities are new, and none of them are AI-specific by definition. What's changing is the scope of workload they have to cover.
The demand-side signal is hard to miss. The FinOps Foundation's State of FinOps 2026 report, based on a survey of nearly 1,200 practitioners representing more than $83 billion in annual cloud spend, found that 98% of FinOps teams now manage AI spend, up from 63% in 2025 and just 31% in 2024. AI cost management ranks as the single most desired skillset FinOps teams want to add over the next year, and when practitioners were asked what tooling capability they most wished existed but doesn't yet, the top answer was granular monitoring of AI spend: tokens, LLM requests, and GPU utilization. In three years, AI spend went from a rounding error to something nearly every FinOps team is responsible for, without a tooling category built specifically to handle it.
Gartner's own research calendar points the same direction. A session at the December 2026 Gartner IT Infrastructure, Operations & Cloud Strategies Conference in Tokyo already refers to the research as the "Magic Quadrant for Cloud and AI Financial Management Tools," shortened to CAIFM, and ties that rename to vendors "increasingly evaluated on more than cloud spend," including how well they enable cost management and optimization of AI workloads specifically. That's one confirmation among several, not the whole case, but it matches what practitioners are already reporting on the ground.
Put plainly: the criteria Gartner already uses to grade CFM tools aren't being replaced. They're being asked to cover a category of spend that didn't exist at meaningful scale the last time most of these tools were built.
What does AI cost management actually involve?
Applied to AI workloads specifically, Gartner's four mandatory CFM capabilities translate into a fairly consistent operating loop, and it's one the FinOps Foundation's own FinOps for AI work describes in similar terms.
Measure. Capture usage at the level AI actually bills at: tokens, not instance-hours. Input tokens, output tokens, cached tokens, and reasoning tokens all carry different rates depending on the provider, so measurement has to happen at the request level, not the account level.
Attribute. Map that usage back to an owner, a team, a product feature, a customer, the same accountability goal that showback and chargeback have always served, applied to a workload where tagging often breaks down. Shared GPU clusters and LLM gateways routinely strip or obscure the signal that traditional tagging depends on.
Optimize. Route requests to the right model tier for the task, smaller and cheaper for classification and extraction, larger and more expensive for tasks that actually need the reasoning depth, and use levers like prompt caching and batch processing where latency allows. This is the AI equivalent of rightsizing and commitment management in traditional cloud cost optimization.
Govern. Set budgets, alerts, and runtime guardrails, hard caps on agentic tool-call depth, per-team token budgets, anomaly detection tuned to how AI spend actually spikes. A traditional monthly reporting cycle can miss an AI cost spike by days; agentic workloads can multiply spend within hours.
Prove value. Translate raw spend into a number the business can act on: cost per inference, cost per successful task, cost per customer. This is the AI cost management version of unit economics, and it's the step most FinOps for AI programs still struggle to close, largely because the measurement layer underneath it isn't solid yet.

How big is the AI cost overrun problem, really?
Bigger than most FinOps teams expected going into it. A survey of 500 finance leaders at organizations already spending on AI, commissioned by DoiT and fielded independently by Sapio Research in February 2026, found that AI now accounts for a mean of 17.6% of total technology spend. Yet 79% of respondents had experienced AI-related cost overruns in the past 12 months, and only 15% could calculate AI ROI without significant bottlenecks.
The most counterintuitive finding in that data: organizations that rate their own FinOps practice as very mature or leading-edge posted the highest overrun rate, at 89%. That's not evidence that mature governance fails. It's more likely that these organizations are running larger, more complex AI initiatives and have the visibility to actually detect overruns that less mature teams simply miss. Maturity surfaces the problem. It doesn't make the underlying measurement gap disappear on its own.
That gap is exactly what AI cost management, and by extension the criteria Gartner is expanding to cover, is trying to close.
Cloud bill shouldn't be a mystery
One platform for AI and Cloud optimization.
What should you look for in an AI cost management tool?
Gartner's existing CFM checklist, configurable dashboards, anomaly detection, AI-driven analytics, utilization monitoring, budget controls, resource optimization, and remediation workflows, still applies. Layered on top of it, a handful of AI-specific requirements separate tools that genuinely manage AI cost from tools that added a line item for it after the fact.
Token- and GPU-level granularity. Account-level or even service-level reporting isn't enough. You need visibility into input, output, cached, and reasoning tokens, split by model and provider, not just a lump-sum API bill.
Multi-provider attribution. Most enterprises run more than one model provider, Anthropic, OpenAI, Google Gemini, AWS Bedrock, often at the same time. A tool that only covers one provider's usage dashboard creates the same kind of tool sprawl that multi-cloud FinOps spent a decade solving for infrastructure.
Coverage for shared and agentic infrastructure. Ask specifically how a tool attributes cost when a request passes through an LLM gateway, or when an agent spawns sub-agents that trigger costs outside the original call. If the answer depends entirely on tags or SDK instrumentation, ask what happens to the spend that infrastructure change breaks before engineering catches up.
Unit economics, not just spend totals. Cost per inference, cost per customer, cost per feature. A dashboard that only shows what was spent, without a path to what that spend produced, is solving half the problem.
Low-lift attribution. Instrumentation that depends on every team tagging every request correctly, forever, tends to degrade the moment usage patterns shift. Approaches that measure consumption at the infrastructure level, rather than depending on tags surviving the trip from request to bill, hold up better as AI architecture keeps changing faster than most engineering teams can document it.
How does this play out across model providers?
The provider-specific pricing guides on DoiT's blog make the same argument from different angles, which is the point: this isn't an Anthropic problem, or a Bedrock problem, or an OpenAI problem. It's the same measurement gap showing up on three different billing structures.
Anthropic's Claude models price input and output tokens separately by tier (Haiku, Sonnet, Opus), so a single agentic workflow can shift cost significantly depending on which tier handles which step, whether prompt caching is in play, and whether requests route through a shared gateway. Amazon Bedrock adds a second axis on top of that: on-demand, provisioned throughput, and batch inference carry different cost and latency trade-offs, and model selection alone can shift per-token rates by 10x to 20x within the same model family. OpenAI's usage and cost integration exists for the same reason: the attribution problem doesn't go away just because the provider changes.
Most enterprises run more than one of these at once, often inside the same application, which is exactly why multi-provider attribution made the checklist above. A CFM tool that only covers one provider's usage dashboard doesn't close the blind spot, it just relocates it. DoiT works across all three, through its Anthropic and OpenAI practices and its native Bedrock, Vertex AI, and Azure AI integrations, on the same problem: making cross-provider AI spend visible, attributable, and defensible as it scales, rather than three separate line items finance discovers after the invoice arrives.
Frequently asked questions
What is AI cost management?
AI cost management is the practice of measuring, attributing, forecasting, and optimizing spend on AI and LLM workloads, tokens, GPU compute, inference, and the infrastructure supporting them. It extends the discipline FinOps built for cloud infrastructure into a category of spend that behaves differently: less stable, more shared, and faster-changing.
What's the difference between AI cost management and AI FinOps?
In practice, the terms are used interchangeably. AI FinOps tends to emphasize the cross-functional operating model, finance and engineering working from shared numbers, while AI cost management more often describes the underlying practice and tooling. Neither term has a single standardized definition yet, which is itself a sign of how new the category is.
Is Gartner adding AI-specific criteria to the Cloud Financial Management Magic Quadrant?
Signs point that way. A session at Gartner's December 2026 conference already refers to the next edition as the Magic Quadrant for Cloud and AI Financial Management (CAIFM) Tools, and the session description ties that rename to vendors being evaluated on how well they manage and optimize AI workload costs, not just traditional cloud spend.
How is AI cost optimization different from cloud cost optimization?
Cloud cost optimization typically means rightsizing, commitment management, and eliminating idle resources on infrastructure with a relatively stable owner and shape. AI cost optimization adds model routing (matching task complexity to the cheapest model that can handle it), prompt caching, and batch processing to that toolkit, and has to account for spend that can shift ownership mid-request in ways traditional infrastructure rarely does.
How common are AI cost overruns?
Common enough to be closer to the norm than the exception. A February 2026 survey of 500 finance leaders, commissioned by DoiT and fielded independently by Sapio Research, found that 79% had experienced AI-related cost overruns in the past 12 months, a rate that climbed to 89% among organizations that rate their own FinOps practice as most mature.
What metrics should an AI cost management program track?
Most programs track some combination of cost per token, cost per inference, cost per feature, and cost per customer or account, alongside the input, output, cached, and reasoning token splits that vary by provider. Engineering teams tend to want workload-level detail; finance teams tend to want account- or feature-level rollups tied to revenue or renewal data.