AI Bills Are Cloud Bills Now: A Practical Guide to LLM Cost Allocation

LLM cost allocation for engineering teams: split OpenAI, Anthropic, Bedrock and Azure OpenAI spend by feature and model, and see it next to your cloud bill.

Xplorr team

The people who build Xplorr

7 min read
AI Bills Are Cloud Bills Now: A Practical Guide to LLM Cost Allocation
In this post
  1. Why is LLM spend hard to allocate?
  2. Step 1: Give every feature its own billing boundary
  3. Step 2: Read cost by model, not only by feature
  4. Step 3: Turn it into a unit cost
  5. Step 4: Put it next to the rest of the cloud bill
  6. Common mistakes
  7. How Xplorr shows AI spend

A year ago the AI line on most engineering budgets was a rounding error. Now it is often the fastest growing line, and it lives in pieces: an OpenAI invoice, an Anthropic invoice, a Bedrock charge inside the AWS bill, an Azure OpenAI charge inside the Azure one.

Finance sees several bills. Engineering sees none of them per feature. That is the gap LLM cost allocation closes: turning “we spent this much on AI” into “this feature cost this much, on this model, for this many requests”. The useful question is not how much you spend on AI, but which feature it is for and whether it is worth it.

Infographic titled What does each AI feature cost. A table of the billing boundary to use per provider: OpenAI projects, Anthropic workspaces, Amazon Bedrock application inference profiles with cost allocation tags, and Azure OpenAI deployments or resources, with what Xplorr collects from each (daily cost, tokens and requests for OpenAI; cost and tokens for Anthropic; cost without token counts from the AWS and Azure bills). A four step flow: split by feature, read by model, compute a unit cost (example: $1,200 a month over 60,000 summaries is 2 cents each), and add the cloud side.
One billing boundary per feature, per provider.

Why is LLM spend hard to allocate?

Three things make it harder than ordinary cloud cost:

  1. It is split across providers with different billing models. OpenAI and Anthropic bill directly by tokens. Bedrock and Azure OpenAI bill through the cloud provider, where the charge sits next to your compute and storage.
  2. The unit is the token, not the resource. There is no instance to tag. One API key can serve five features, and the invoice will not tell them apart.
  3. Prices differ by model, by a lot. A feature that moves from a small model to a large one can multiply its cost without any change in traffic. Cost per feature has to be read together with cost per model.

Step 1: Give every feature its own billing boundary

Allocation starts before the bill. Each provider has a unit that its usage and cost reports can split by, and the simplest allocation is one of those units per feature (and, ideally, per environment):

  • OpenAI: create a project per feature or product area. Usage and cost reports break down by project, and each project has its own API keys.
  • Anthropic: create a workspace per feature or team. Usage and cost reports break down by workspace, and keys belong to a workspace.
  • Amazon Bedrock: create an application inference profile per feature and attach cost allocation tags to it. Calls made through that profile carry the tags into Cost Explorer and the Cost and Usage Report.
  • Azure OpenAI: use a separate deployment or resource per feature, tagged with the feature and team, so the cost splits in Azure Cost Management.

If a single shared key serves everything today, splitting it is the most valuable hour you can spend on AI cost. Nothing downstream can recover a split the provider never recorded.

Step 2: Read cost by model, not only by feature

Once spend splits by feature, add the model dimension. Most providers report usage per model per day, with input and output tokens counted separately (output tokens usually cost several times more than input). The patterns worth looking for:

  • A feature whose cost rose while its request count did not. Usually a model change, a longer prompt or a bigger context window.
  • A feature on a large model doing a job a small one could do, such as classification, extraction or routing.
  • Output heavy features, where limiting response length saves more than any prompt trimming.

Step 3: Turn it into a unit cost

A total per feature is a start. A unit cost is what lets someone decide whether the feature is worth it: AI cost per request, per active user, or per document processed. Divide the feature’s monthly AI spend by the count your product already tracks. As an example, a support summary feature that costs $1,200 a month and runs 60,000 times costs 2 cents a summary. That number can be compared with the value of the summary, which the total cannot.

Step 4: Put it next to the rest of the cloud bill

AI features rarely cost only tokens. They run on compute, read from vector stores and databases, and move data. Allocating the model bill on its own gives half the answer. The full cost of a feature is its LLM spend plus the cloud resources behind it, which is why AI spend belongs in the same view, with the same alerts, as the rest of your cloud cost. A runaway retry loop against a model API is a cost anomaly like any other, and it should reach the same channel.

Common mistakes

  • One key for everything. It works until the first time someone asks what a feature costs.
  • Monthly review only. Token spend can climb quickly after a release. Daily data and an alert catch a prompt or retry bug the same week.
  • Ignoring the cloud side. Bedrock and Azure OpenAI charges are easy to miss because they sit inside AWS and Azure bills, not on a separate AI invoice.

How Xplorr shows AI spend

Xplorr pulls OpenAI and Anthropic spend next to AWS, Azure and GCP. With an OpenAI organization Admin API key, it collects daily cost, tokens and requests per project and model. With an Anthropic Admin key, it collects daily cost and tokens per workspace and model. Bedrock and Azure OpenAI appear from the cloud bill itself (cost, without token counts). All of it sits in the same console, and anomaly detection runs on AI spend as it does on cloud services, with alerts in Slack and by email.

See AI spend for the view, OpenAI and Anthropic cost tracking for how it compares with the providers’ own dashboards, and cost sources for other spend outside the cloud bill. The AI spend guide covers the keys each provider needs. Xplorr is free during the private beta.

Keep reading

See how Xplorr helps → Features

ShareLinkedInX

Written by

Xplorr team

The people who build Xplorr

Written together by the engineers who build Xplorr: the AWS, Azure, GCP and Kubernetes collectors, the console, and the alerting behind them.

About Xplorr

Related posts

All articles
Inside the Xplorr Console: A Walkthrough of the Demo Workspace
Product

Inside the Xplorr Console: A Walkthrough of the Demo Workspace

Eight screens, from connecting an account to reading a Kubernetes namespace bill, illustrated with our seeded demo workspace. Every figure shown is synthetic sample data for a fictional company, not a customer result.

Xplorr team

9 min read

What Building Cloud Cost Integrations Actually Teaches You
Engineering

What Building Cloud Cost Integrations Actually Teaches You

Five findings from the AWS, Azure, GCP and OpenCost collectors: a placeholder that hides unused storage, an API with no step parameter, and cost data that is absent unless you opt in.

Xplorr team

10 min read

Free during the private beta

See your AWS, Azure and GCP spend in one place

Connect a cloud account and find what is driving the bill. Every feature is free while Xplorr is in private beta.