AI model costs look simple on the surface. ARK Invest’s analysis of Artificial Analysis benchmark data found inference costs for capable models declining 95% at an annual rate, yet the bill keeps growing. Between input tokens, output tokens, cached pricing, fine-tuning fees, and the infrastructure to run it all, the per-token price is only the most visible line on the bill. If you don’t break that bill apart, you can’t govern it.
This guide breaks down every component of AI model pricing, compares costs across OpenAI, Anthropic, Google, and self-hosted options, and walks through the strategies that actually reduce spend without sacrificing capability.
Knowing what a token costs is only useful if you can see where your tokens are actually going. The platforms below ingest spend from multiple AI providers, and they’re evaluated on provider coverage, allocation depth, anomaly detection, and whether AI spend sits unified with cloud spend or bolted on separately. Finout, an enterprise-grade FinOps platform for cloud and AI spend, leads the list for teams that need AI and cloud costs allocated under one standard.
Scope: platforms that track multi-provider AI and LLM spend, as described on each vendor’s official pages in September 2026.
| Platform | Multi-Provider Ingestion | Automated Allocation | Anomaly Detection | Unified with Cloud Spend | Best For |
|---|---|---|---|---|---|
| Finout | OpenAI, Anthropic, Cursor, Bedrock, Vertex AI, Azure OpenAI, and more | AI-Powered VTags, no code changes | Yes, with Detection and Investigation Agents | Yes, in the MegaBill | Teams that want AI and cloud cost under one FinOps standard |
| Amnic | OpenAI, Anthropic, Gemini, Bedrock | Token attribution to teams, users, and cost centers | Yes, threshold and budget alerts | Teams that want AI token and cloud cost views in one platform | |
| Holori | OpenAI, Anthropic, Bedrock, Vertex AI, Azure OpenAI, LiteLLM | Virtual Tags, cost centers | Yes, spike, token, and budget alerts | Yes | Teams that want a multi-cloud FinOps view that includes LLM spend |
| Datadog Cloud Cost Management | OpenAI, Anthropic, Bedrock, Gemini, Vertex AI, GitHub Copilot, Cursor | Attribution by user, project, or model | Yes | Yes | Teams that want AI spend inside Datadog Cloud Cost Management |
| Harness Cloud & AI Cost Management | OpenAI, Anthropic, Bedrock, Vertex AI | Chargeback and showback by team, model, and product | Yes, across cloud, LLM, and GPU workloads | Yes | Teams that want AI chargeback and showback next to cloud spend |
| Braintrust | LLM calls via SDK or gateway (OpenAI, Anthropic, Google, and more) | Code-level trace tags | Cost-threshold alerts | Teams tracking LLM cost next to evaluation quality |
Finout is an enterprise-grade FinOps platform for cloud and AI spend, built for the agentic era. FinOps for AI ingests OpenAI, Anthropic, Cursor, Amazon Bedrock, Google Vertex AI, and Azure OpenAI spend, AI-Powered VTags map it to teams and features automatically, Billy answers cost questions in plain English, and the Detection and Investigation Agents flag anomalies with root-cause context. All of it sits in the MegaBill alongside cloud, Kubernetes, and SaaS costs.
Best for: Teams that want AI and cloud cost under one FinOps standard.
Key consideration: Pricing is a flat fee based on tiers of committed spend, detailed on the Finout pricing page.
Amnic is a FinOps platform whose AI Token Management module tracks token usage and cost from LLM providers including OpenAI, Anthropic, Gemini, and Amazon Bedrock in a unified view with cloud spend, and alerts on token thresholds, budget limits, and anomalous cost spikes.
Best for: Teams that want AI token and cloud cost views in one platform.
Key consideration: AI token attribution covers teams, users, and cost centers.
Holori is a multi-cloud FinOps platform that shows spend from OpenAI, Anthropic, AWS Bedrock, Vertex AI, Azure OpenAI, and LiteLLM in one normalized view alongside AWS, Azure, GCP, and OCI costs, with Virtual Tag allocation and alerts for cost spikes, unusual token consumption, and budget overruns.
Best for: Teams that want a multi-cloud FinOps view that includes LLM spend.
Key consideration: LLM coverage runs through six named sources: OpenAI, Anthropic, Bedrock, Vertex AI, Azure OpenAI, and LiteLLM.
Datadog Cloud Cost Management analyzes AI spend across Amazon Bedrock, Anthropic, Google Gemini, OpenAI, Vertex AI, GitHub Copilot, and Cursor, shows it alongside cloud infrastructure costs, and tracks cost anomalies.
Best for: Teams that want AI spend inside Datadog Cloud Cost Management.
Key consideration: AI cost views sit inside Datadog’s broader Cloud Cost Management product.
Harness Cloud & AI Cost Management captures spend from OpenAI, Anthropic, AWS Bedrock, and Google Cloud Vertex AI, with chargeback and showback by team, model, and product and anomaly detection across cloud, LLM, and GPU workloads.
Best for: Teams that want AI chargeback and showback next to cloud spend.
Key consideration: AI provider coverage on the product page names four sources: OpenAI, Anthropic, AWS Bedrock, and Vertex AI.
Braintrust is an agent observability and evaluation platform that tracks latency, cost, and quality for traced LLM calls in real time, with cost-threshold alerts.
Best for: Teams tracking LLM cost next to evaluation quality.
Key consideration: Cost figures are estimates attached to each traced LLM call.
AI models charge on a per-token consumption model, where a token equals roughly three-quarters of a word. Your costs depend on whether tokens are input (the prompts and context you send) or output (the content the model generates back). Output tokens cost more than input tokens on every major provider’s price list because generation requires more compute, as the provider comparison below shows.
A cost breakdown separates your AI bill into distinct categories, showing exactly where spend originates. Think of it like itemizing a restaurant bill instead of just seeing the total. Once you can see the line items, you can start asking better questions about what’s worth the money.
The question usually starts with finance or leadership: what does this model, feature, or customer cost us, and is it worth it? AI spend makes that hard to answer for three reasons:
The real challenge is financial accountability for AI spend. When multiple teams share API keys or when AI features are embedded across different products, no one owns the cost. And when no one owns it, no one optimizes it. The problem gets worse because AI spend sits in four places at once: cloud AI services such as Amazon Bedrock and Vertex AI inside your cloud bill, direct contracts with providers like OpenAI and Anthropic, token-based developer tools such as Cursor and GitHub Copilot, and SaaS tools that are moving to usage-based pricing. Without a common control plane that consolidates these into a single view, FinOps teams are left reconciling spreadsheets instead of driving action.
Understanding what you’re actually paying for is the first step toward controlling AI spend. Not every provider charges for all of the components below, but each one can show up on your bill. For a deeper look at per-token rates, see our guide to what AI token cost includes.
Input tokens are the text you send to the model: your prompts, system instructions, and any context you include. Providers charge per million input tokens, with rates varying by model tier. Within a single provider’s lineup, the input rate for a budget model can be a small fraction of the flagship rate. The provider comparison below lists current rates.
Output tokens are what the model generates in response. Generation requires more compute than processing input, and output rates reflect that: on current flagship price lists, output runs about five to six times the input rate. Reasoning models add another layer, since the thinking tokens they generate are billed as output even though you never see them.
Some providers offer discounted pricing when the same prompt prefix is reused across requests. On Google’s Gemini API price list, for example, cached input for Gemini 3.8 Flash costs $0.075 per million tokens against a standard input rate of $0.75, a 90% discount. If your application sends repetitive queries, caching can meaningfully reduce spend. Cache hit rate decides how much of your input gets recomputed at full price, and that rate flows straight into cost per request, cost per customer, and the margin on any AI feature you charge for.
The context window is the maximum tokens a model can process in a single request. Larger context windows cost more to use. A 128K context window is powerful, but sending 100K tokens when 10K would suffice wastes money.
Fine-tuning trains a model on your own data to improve performance for specific tasks. This involves upfront training costs plus ongoing inference costs for the custom model. Availability varies by provider, and OpenAI is winding down its fine-tuning platform (see below). Check whether the option still exists before you plan for it.
If you self-host open-source models like Llama or Mistral, you pay for GPU compute, storage, and orchestration instead of per-token API fees. This shifts costs from variable to fixed, which can work well at scale but requires engineering investment. Beyond GPU hours, you’re also paying for model serving infrastructure, load balancing, monitoring, and the MLOps team to keep it all running. The total cost of compute for self-hosted models often surprises teams who focus only on the hardware line item.
| Component | What It Covers | OpenAI | Anthropic | Self-Hosted | |
|---|---|---|---|---|---|
| Input tokens | Prompts and context | ✓ | ✓ | ✓ | N/A |
| Output tokens | Generated responses | ✓ | ✓ | ✓ | N/A |
| Cached tokens | Reused prompt prefixes | ✓ | ✓ | ✓ | N/A |
| Fine-tuning | Custom model training | Winding down | Varies | Varies | ✓ |
| Infrastructure | GPU compute and storage | N/A | N/A | N/A | ✓ |
Prices below are list prices for standard API tiers as of September 2026. Providers change lineups often. Check the provider pages before you budget.
OpenAI pricing now centers on the GPT-6 family. On OpenAI’s API pricing page, GPT-6 Astra costs $10 per million input tokens and $50 per million output tokens, GPT-6 Sol costs $2 and $10, and GPT-6 Luna costs $0.10 and $0.50, all at the Standard tier with short context. Cached input costs a tenth of the standard input rate, the Batch tier halves both input and output rates, and long-context requests are priced higher. OpenAI also says it is winding down its fine-tuning platform, which is no longer open to new users. GPT-4o, o1, and o1-pro no longer appear on the current pricing page, and with a catalog that changes this often, choosing the right tier for each task matters more than memorizing a price.
Anthropic’s Claude API pricing follows a similar tiered structure. According to Anthropic’s pricing page, Claude Opus 5.5, its daily driver for agentic coding and enterprise work, costs $4 per million input tokens and $20 per million output tokens. Claude Sonnet 5 costs $2 and $10, Claude Haiku 4.5 costs $1 and $5, and Claude Fable 5.1, built for long-running agents, costs $10 and $50. Prompt-cache reads cost $0.20 per million tokens on Opus 5.5 and Sonnet 5, and batch processing saves 50%.
Google’s Gemini models are available through the Gemini API and Vertex AI. On Google’s Gemini API pricing page, Gemini 3.1 Pro Preview costs $2 per million input tokens and $12 per million output tokens for prompts up to 200K tokens, rising to $4 and $18 above that. Gemini 3.8 Flash costs $0.75 and $3.75 through December 31, 2026, doubling to $1.50 and $7.50 on January 1, 2027, and Gemini 3.5 Flash-Lite costs $0.30 and $2.50. Output prices include thinking tokens, and the Batch API cuts costs by 50%. Gemini pricing can also appear in different billing contexts depending on how you access the models.
Open-weight models such as Llama and Mistral eliminate per-token API fees entirely. Instead, you pay for GPU instances by the hour whether or not they’re busy, plus serving, monitoring, and the people who run it. The break-even point depends on your volume and operational capacity.
| Provider | Model Tiers | Pricing Structure | Key Pricing Note |
|---|---|---|---|
| OpenAI | GPT-6 Astra, GPT-6 Sol, GPT-6 Luna | Per-token, tiered by capability | Batch tier at half the Standard rate; cached input at a tenth of standard |
| Anthropic | Opus 5.5, Sonnet 5, Haiku 4.5, Fable 5.1 | Per-token, tiered by capability | Batch processing saves 50%; discounted prompt-cache reads |
| Gemini 3.1 Pro Preview, 3.8 Flash, 3.5 Flash-Lite | Per-token, higher rates above 200K-token prompts on Pro | Batch API cuts costs 50%; Flash rates double on January 1, 2027 | |
| Self-hosted | Llama, Mistral | Compute-based (GPU hours) | No per-token fees; you pay for idle GPU time too |
A $0.10-per-million-token model that needs three attempts to produce a usable answer still looks cheap on the invoice, but every failed attempt adds latency, retry logic, and review time that never shows up in the per-token rate. The right choice depends on the task, and getting it wrong at scale is one of the fastest ways to inflate AI spend.
When evaluating models, consider four dimensions:
| Use Case | Recommended Tier | Why |
|---|---|---|
| Simple queries, classification | Budget (GPT-6 Luna, Claude Haiku 4.5, Gemini 3.5 Flash-Lite) | Low complexity doesn’t justify premium pricing |
| Code generation, analysis | Mid-tier (GPT-6 Sol, Claude Sonnet 5, Gemini 3.8 Flash) | Requires reasoning but not maximum capability |
| Complex reasoning, research | Flagship (GPT-6 Astra, Claude Opus 5.5, Gemini 3.1 Pro Preview) | Quality matters more than cost per token |
If you’re only tracking what the API charges per million tokens, you’re missing the retries, the storage, the evaluation runs, and the guardrails that quietly inflate your real cost. These hidden line items are where AI’s financial risk actually lives.
When a request times out or returns an unusable answer, many applications retry automatically, and every attempt the model processes is billed. Rate limits add a second cost when fallback logic sends overflow traffic to a different model with a different rate.
Moving data between cloud regions or storing conversation history and embeddings adds incremental costs. If your AI application stores every interaction for fine-tuning or compliance, storage costs compound over time.
Training runs, evaluation datasets, and iterative tuning all consume billable compute before you reach production. The bill for a tuning job scales with dataset size, training time, and how many iterations it takes to reach production quality.
Monitoring, logging, and safety layers add costs on top of base model pricing. Content moderation APIs, guardrail services, and evaluation frameworks all have their own billing meters.
Understanding your unit economics requires connecting spend data to usage metrics.
Start by consolidating invoices from OpenAI, Anthropic, and any other providers into a single view. When teams use separate accounts or API keys, spend fragments across billing contexts. Finout ingests AI provider costs automatically alongside cloud spend into the MegaBill, and with Billy, Finout’s AI FinOps assistant, you can ask natural-language questions like “What did Team A spend on OpenAI last month?” and get instant, chart-backed answers without building custom queries.
Pull usage metrics from provider dashboards or API logs. Track input and output tokens separately since they have different costs and different optimization levers.
Divide total spend by tokens, API calls, or active users to get unit costs. If your chatbot feature costs $500/month and serves 10,000 users, your cost per user is $0.05.
Tag or allocate costs to business dimensions to answer questions like “How much does Team A spend on AI?” Virtual tagging can map untagged AI spend to the right owner without code changes. AI-Powered VTags take this further, scanning names, labels, and metadata across connected cloud, Kubernetes, SaaS, and AI sources to propose hundreds of allocation rules automatically. You approve, edit, or reject them in bulk, and the rules apply retroactively. For teams building internal tooling, Finout’s Allocation API supports “allocation as code,” making it possible to export fully allocated AI spend to your own analytics or billing systems.
Allocation assigns shared AI costs to specific teams, products, or customers, and it’s what turns a provider invoice into showback, unit economics, and a forecast someone owns. It’s harder than traditional cloud allocation for two reasons:
Three approaches cover most cases:
Normalization is the step most teams skip. Every provider structures billing and names models differently, and the same model can be bought through more than one billing path. Finout’s Canonical AI Taxonomy applies predefined Virtual Tags to LLM spend, including a Model Channel tag that identifies the billing path, such as Bedrock, Anthropic, OpenAI, or Vertex AI, and a Token Type tag that splits cost lines into input, output, cache read, and cache write.
AI usage is harder to predict than traditional compute because it depends on user behavior, prompt complexity, and feature adoption.
Finout’s Financial Plans move these forecasting strategies from spreadsheets into a governed environment. You can set budgets by team, feature, or segment, sync actuals against plan in real time, and get alerted when projected spend exceeds budget. For multi-year AI roadmaps, custom lines let you reserve space for future expenses, like a planned model migration or a new agentic workflow, keeping your financial plan connected to how engineering actually works.
Model routing uses lightweight models for simple tasks and reserves flagship models for complex reasoning. A classification task rarely needs a flagship model, and a budget tier can handle it at a small fraction of the per-token rate.
Semantic caching stores responses for repeated or similar queries. If 20% of your queries are near-duplicates, caching eliminates 20% of token consumption.
Every unnecessary token costs money. Remove redundant instructions, summarize long inputs, and avoid filling the context window when a smaller context would suffice.
If a workload doesn’t need an answer in real time, send it through a batch tier. OpenAI, Anthropic, and Google all price batch requests at roughly half their standard rates, as the provider comparison above shows.
Configure alerts that fire when AI spend exceeds thresholds. A single misconfigured loop can keep running up charges overnight before anyone looks at the bill. Finout’s Detection Agent continuously scans your AI, cloud, and SaaS environments for waste, drift, and cost anomalies, surfacing only financially relevant findings. When it flags something, the Investigation Agent runs root-cause analysis, mapping the anomaly to its blast radius, ownership, and history, giving your team context to act on instead of raw alerts.
Batch, cached, and long-context rates now appear as separate columns on OpenAI’s, Anthropic’s, and Google’s price lists, which means the same model can carry several prices depending on how you call it.
AI agents that chain multiple model calls dramatically increase token consumption compared to single-turn queries. An agent making 10 model calls costs roughly 10 times a simple query, and BCG’s AI Radar 2026 found CEOs have committed more than 30% of their organizations’ AI investment to agentic AI this year. Gartner predicts that AI inference costs per agentic workflow will increase more than fivefold through 2028, even as tokens get cheaper, a dynamic Gartner calls the Inference Paradox.
To govern this spend, Finout’s MCP server lets AI agents and developer tools query cost data directly. An engineering copilot can use it to answer “did my PR change spend?”, and an incident agent can route cost anomalies to the right owner, all within the same governed data layer that powers your FinOps workflows.
Routing between models based on task complexity is becoming a core design choice rather than an optimization afterthought. Gartner’s Inference Paradox analysis argues that inference tiering, routing, and orchestration are how product teams protect margins as agentic workloads grow.
Managing AI costs alongside cloud spend requires one platform, not four billing exports. FinOps for AI ingests OpenAI, Anthropic, Cursor, Amazon Bedrock, Google Vertex AI, Azure OpenAI, GitHub Copilot, and fal.ai spend. Billy gives you natural-language answers to cost questions. FinOps Agents handle detection, investigation, and governed orchestration across your environment. The MCP server gives your internal agents programmatic access to the same governed cost data.
All of it lands in the MegaBill, which brings AWS, Azure, GCP, OCI, Kubernetes, SaaS, and AI costs into one unified view for allocation, budgeting, anomaly detection, and optimization.
Lyft used Finout to improve cost attribution from 80% to more than 96% and to cut issue detection from weeks to days.
The result: one FinOps standard for cloud and AI spend, built for the agentic era.