Every AI Cost Conversation Has an Altitude
The numbers become harder to obtain as they become more useful to the business.
Every AI cost conversation has an altitude.
At the bottom, the numbers are easy to collect. At the top, they are useful for decisions. Most organizations can see the bottom clearly while speaking as though they manage the top.
That distance is the central management problem.
The six rungs
Picture AI cost as a staircase:
Cost per token
Cost per call
Cost per task attempted
Cost per task completed successfully
Cost per business outcome
Cost per dollar of value
The first two rungs arrive with the service. The model provider meters tokens and calls because those are the units it sells.
The middle rungs belong to the workflow. To measure them, the organization has to define a unit of work, record attempts, identify completions, and separate success from failure.
The top rungs belong to the business. They require a named outcome, a defensible value, and agreement about where that value lands in the P&L.
One number does the most work across this series: the fully loaded cost per successful outcome. Fully loaded means every cost the workflow consumes, not just the model bill: model and platform, data, evaluation, human review, and operating support. Successful means the outcome met an agreed definition, not merely that a call returned. Later articles use this number without redefining it.
The visibility inversion
The staircase contains an uncomfortable inversion. It is the asymmetry from the first article drawn as a picture: the seller meters the bottom, and the buyer has to construct the top.
Cost is clearest at the bottom. The invoice can report consumption to several decimal places. Value is almost invisible there.
As you climb, value becomes clearer. You can see whether a ticket stayed closed, whether a payment settled, or whether a claim avoided further handling. But measurement becomes more expensive and more contested.
The two things an executive needs to make a decision, what it cost and whether it was worth it, are clearest at opposite ends.
That is why so many AI dashboards are busy and unsatisfying. They report the bottom of the staircase in high resolution and leave the top to narrative.
Measured altitude and claimed altitude
A useful review places the organization on the staircase twice.
The measured altitude is the highest rung supported by a real, attributable number. If the organization knows tokens and calls by team but cannot count completed units of work, its measured altitude is Rung 2.
The claimed altitude is the rung implied by the business case. If a board presentation claims customer-service savings, improved retention, or margin growth, the argument is being made at Rung 5 or 6.
The gap between those positions is the diagnosis.
For example, an organization may operate a customer-service copilot with detailed model and token reporting. The business case claims lower cost to serve. Yet the company does not measure cases resolved, repeat contacts, escalation rates, or human handling time. It measures at Consumption and claims at Value.
The problem is not imperfect data. The problem is an unacknowledged gap between the evidence and the decision.
How to climb without building a measurement empire
Climbing the staircase does not require instrumenting every AI workflow at once.
Start with one material workflow and move one rung higher:
If you know tokens, define the unit of work.
If you know attempts, count completions.
If you know completions, define and verify success.
If you know successful outcomes, connect them to an accepted business effect.
Each move answers a question the previous rung could not answer. Each also exposes which assumption was carrying the business case.
The goal is not to abandon the lower rungs. They remain necessary for billing, engineering, and forecasting. The goal is to stop asking them to stand in for the upper ones.
The executive question
Before adding another AI metric, ask:
Which rung are we actually managing, and which rung are we claiming to manage?
The distance between those answers is where the next measurement investment belongs.
A question for readers
What is the highest rung your organization can support today with a real number?
Onward,
Raja
Raja Pabba is the founder of CloudMetrics and writes The CAIO Review on enterprise AI operating discipline. Subscribe at caioreview.com.



