On this page
Put Coworker to work on your stack.
Connect Salesforce, Slack, Jira and run your first agent in minutes.
Enterprise AI
AI Cost Optimization: Where the Money Goes and 9 Levers That Cut It
Coworker AI's guide to AI cost optimization: where token and seat spend goes, 9 levers from model routing to prompt caching, and how to measure cost per task.
AI cost optimization is the practice of lowering what your organization pays for AI without lowering the quality of the work the AI does. The unit that matters is cost per completed task, and you reduce it by sending each task to the cheapest model that passes your quality bar, sending less context, caching what repeats, batching what can wait, capping output and agent loops, and governing seats.
Almost every FinOps team now owns some version of this job. In the FinOps Foundation's State of FinOps 2026 survey of 1,192 respondents, 98% said they manage AI spend, up from 31% two years earlier. In Zylo's 2026 SaaS Management Index, 78% of the 218 IT leaders surveyed reported unexpected charges tied to consumption-based or AI pricing in the previous 12 months.
I wrote this guide for CIOs and IT and AI leaders whose AI line has two parts, seats and tokens, and both are climbing. I checked every price, discount and multiplier below against the vendor's own documentation on October 6, 2026. Prices change often, so treat them as a dated snapshot. If you are still choosing a platform, start with the enterprise AI pricing comparison. This guide is about lowering the bill you already have.
How is AI cost optimization different from cloud cost optimization?
The arithmetic is the same. The FinOps Foundation's FinOps for AI overview notes that the basic price times quantity equation still applies: you lower cost by paying a lower rate or by using less. What changes is who sets the quantity, and how fast it moves.
In cloud, quantity follows provisioning. Someone launches an instance and it runs until someone stops it. In AI, quantity follows design: the prompt, the retrieval step, the number of tools exposed and the number of turns an agent takes decide how many tokens a request consumes. Microsoft's guide to optimizing AI workload costs on Azure gives two examples. The same user can cost $0.001 one minute and $0.40 the next, depending on context length, retrieval depth and the model the request was routed to. And a naive retrieval-augmented generation prompt can grow from 800 to 12,000 tokens after a single product change.
The people creating the spend have changed too. The FinOps Foundation points out that what it calls non-traditional groups, such as product, marketing, sales and leadership, now contribute to AI costs directly. Its framework has a FinOps for AI section that suggests treating important AI spend as its own scope, with shorter forecasting windows and funding that gets revisited more often until forecasts improve.
AI cost management, AI FinOps and LLM cost optimization compared
These terms overlap, and vendors use them loosely. Here is how I use them in this guide.
| Term | What it covers | Who usually owns it | Typical tooling |
|---|---|---|---|
| AI cost management | Visibility, allocation, budgets and governance across all AI spend: seats, APIs and cloud | IT, FinOps, procurement | FinOps platforms, SaaS management, provider consoles |
| AI FinOps (FinOps for AI) | The FinOps operating model applied to AI: shared accountability, showback and chargeback, unit metrics, forecasting | FinOps team with engineering and finance | Billing data in a common format such as FOCUS, FinOps platforms |
| LLM cost optimization | Request-level engineering: model choice, context size, caching, batching, output limits | Platform and AI engineering | Gateways, observability tools, provider features |
| AI cost optimization | All of the above, aimed at a lower cost per unit of business value | CIO or CTO, with finance | All of the above |
AI cost management tells you where the money goes. LLM cost optimization changes how much each request uses. You need both, because a dashboard that shows spend does not shrink a prompt.
Where does AI spend actually go?
Spend splits into seats, tokens and the infrastructure that serves them. The clearest market-level split I found is in Menlo Ventures' 2025 State of Generative AI in the Enterprise, published December 9, 2025. Menlo estimates companies spent $37 billion on generative AI in 2025, up from $11.5 billion in 2024. Of that, $19 billion went to applications and $18 billion to infrastructure, including $12.5 billion on foundation model APIs. Menlo's sizing excludes chips, inference and model serving on the cloud platforms, and AI features built into existing software, so the real total is larger.
Inside one company, the spend is scattered. The Tokenomics Foundation's State of Tokenomics report, released September 23, 2026 with 472 responses across 11 industries, found that 96% of respondents use model providers such as Anthropic and OpenAI directly, 87% use cloud token providers such as AWS Bedrock, Google Vertex and Azure Foundry, and 64% use AI embedded in tools such as Cursor or Databricks Genie. Half spread their AI consumption across four of the six procurement and hosting channels the survey tracked.
The budgets are growing fast. CloudZero's State of AI Costs survey of 500 U.S. software engineers at manager level and above, run in March 2025, put average monthly AI spend at $62,964 in 2024, and its findings suggested a rise to $85,521 in 2025, a 36% increase. The share of organizations planning to spend more than $100,000 a month was set to rise from 20% to 45%.
Seats: the fixed line that now has a meter
Seat licenses look predictable, and that is changing. On October 6, 2026, Anthropic's Claude pricing page listed Claude Team standard seats at $20 per seat per month billed annually ($25 billed monthly) and premium seats at $100 ($125 monthly) with 5x the usage of a standard seat. Claude Enterprise was listed as a seat price plus usage at API rates: $20 per seat per month billed annually, with usage cost that scales with model and task. OpenAI's business pricing page lists the same two seat tiers for ChatGPT Business, standard at $20 and premium at $100 a month billed annually, with the premium seat carrying 5x the usage, and says credit-based and token-based pricing are available on Enterprise plans.
Two consequences follow. A premium seat costs five times a standard one, so tiers should go to the people who actually use the extra capacity. And where a seat includes metered usage, the token levers later in this guide apply to the seat budget too.
Shadow purchasing adds another layer. Zylo's 2026 index, built on more than 40 million SaaS licenses and $75 billion in spend under management, found AI-native application spend up 108% year over year, and up 393% at organizations with more than 10,000 employees. Expense-based SaaS spend rose 267%, with ChatGPT now the most expensed application, and, measured against recommended utilization levels, organizations left an average of 36% of their SaaS licenses unused. That last figure covers all SaaS, not AI alone, but it is a useful benchmark to hold AI seats against. For per-user benchmarks, see how much enterprise AI should cost per user and the ChatGPT Enterprise pricing breakdown.
Tokens: the variable line
All seven current models in the table below price output tokens at five times input tokens, and the gap between the most and least expensive model is wide.
| Model (list price, October 6, 2026) | Input per 1M tokens | Cached input (cache read) per 1M | Output per 1M tokens |
|---|---|---|---|
| Claude Fable 5.1 | $10.00 | $0.25 | $50.00 |
| Claude Opus 5.5 | $4.00 | $0.20 | $20.00 |
| Claude Sonnet 5.5 | $2.00 | $0.20 | $10.00 |
| Claude Haiku 4.5 | $1.00 | $0.10 | $5.00 |
| GPT-6 Astra | $10.00 | $1.00 | $50.00 |
| GPT-6.1 Sol | $2.00 | $0.10 | $10.00 |
| GPT-6 Luna | $0.10 | $0.01 | $0.50 |
Sources: Anthropic prompt caching pricing and OpenAI API pricing, standard tier, OpenAI short-context rates. Writing to a cache costs more than reading from it; lever 4 below covers the rules.
The spread is 10x between Anthropic's most and least expensive models in this table, and 100x at OpenAI. That spread is the main reason model routing works.
Reasoning is the hidden part of output. OpenAI's reasoning guide says reasoning tokens are not visible through the API but are billed as output tokens, and that a model may generate anywhere from a few hundred to tens of thousands of them depending on the problem. Anthropic's extended thinking docs report thinking tokens as part of billed output in the same way. OpenAI also warns that a request can hit its output limit mid-reasoning and charge you for input and reasoning with no visible answer.
Coding deserves its own line in the budget. Menlo calls it the largest category across the application layer, at $4.0 billion in 2025 and 55% of departmental AI spend. Anthropic's Claude Code cost guide says enterprise deployments average around $13 per developer per active day and $150 to $250 per developer per month, with 90% of users staying below $30 per active day. The Claude Code pricing guide covers the plan options, and the Enterprise AI Price Index tracks how platform prices move each quarter.
Hidden context: the tokens nobody budgets for
Much of what an agent reads is overhead the person asking never sees. Four sources account for a lot of it.
Tool definitions are billed as input. Anthropic's pricing documentation counts tool names, descriptions and schemas among the input tokens you pay for. Its post on advanced tool use gives a five-server example in which 58 tools consume about 55,000 tokens before the conversation starts, and says Anthropic has seen tool definitions consume 134,000 tokens before optimization.
Intermediate results pass through the model. In its post on code execution with MCP, Anthropic describes an agent moving a meeting transcript from Google Drive into Salesforce. The full transcript flows through the model twice, and for a two-hour sales meeting that could mean processing an additional 50,000 tokens.
Retrieval fans out. Microsoft's Azure guide warns that each chat turn might issue three to eight hidden queries that never show up in your logs.
Agents multiply all of it. In Anthropic's write-up of its multi-agent research system, agents typically used about 4x more tokens than chat interactions, and multi-agent systems about 15x more.
Here is what that looks like on a bill. Anthropic's Claude Code cost guide shows an example session on Claude Sonnet 4.6 that cost $0.55. I priced each line at Anthropic's list rates for that model, assuming 5-minute cache writes, which reproduces the $0.55 total:
| Token type in the example session | Tokens | List rate per 1M tokens | Cost |
|---|---|---|---|
| Cache reads (context reused from earlier turns) | 940,000 | $0.30 | $0.282 |
| Cache writes (5-minute lifetime) | 50,000 | $3.75 | $0.188 |
| Uncached input | 1,200 | $3.00 | $0.004 |
| Output | 5,300 | $15.00 | $0.080 |
| Total | 996,500 | $0.55 |
Context made up more than 99% of the tokens and about 86% of the cost, even with caching doing its job. Billed as ordinary input, the same session would have cost about $3.05. Output is the expensive token, but context is where the volume is. The context window explainer covers how that space fills up.
Pricing modifiers that change the bill
A few line items never appear in a headline price per million tokens.
| Modifier | Anthropic (Claude API) | OpenAI (API) |
|---|---|---|
| Long context | Claude 4.6 and later models bill the full 1M-token window at standard rates | Prompts over 272K input tokens pay 2x the input rate and 1.5x the output rate on the flagship models |
| Data residency | US-only inference costs 1.1x | Regional processing endpoints carry a 10% uplift for models released on or after March 5, 2026 |
| Faster output | Fast mode for Opus 5.5 costs 2x standard pricing, for up to 2.5x faster output | The Fast tier costs 2x standard rates |
| Built-in web search | $10 per 1,000 searches, plus tokens for the results | $10 per 1,000 calls, plus search content tokens at model rates |
Sources: Anthropic pricing documentation, Claude pricing and OpenAI API pricing, checked October 6, 2026.
None of these are wrong to use. Data residency may be a compliance requirement, and fast mode may be worth paying for in a flow where someone is waiting. The point is to choose them per workload instead of by default.
How do you measure AI cost per task?
The FinOps Foundation's FinOps for AI overview suggests KPIs such as cost per inference, cost per token, cost per API call, anomaly detection rate, return on investment and time to business value. All are useful. I would add one more as the headline number: cost per completed task.
Cost per token tells you whether you bought tokens cheaply. Cost per task tells you whether the work was cheap. A model that costs half as much per token but needs three attempts to pass review is more expensive per task, and only the second number shows it.
The survey data says this is where companies get stuck. In the State of Tokenomics report, 43% of respondents named proving value or ROI as their single biggest challenge, while only 7% named cost or pricing complexity. 39% were not confident they could connect AI spend to a measurable business outcome their CFO would accept. The very confident respondents shared two traits: spend metered and attributable across workloads and teams, and concrete business output metrics such as tickets, pull requests and revenue.
The formula
Cost per completed task = (input tokens × input rate) + (cached tokens × cache rate) + (output and reasoning tokens × output rate) + tool and search fees, divided by the number of tasks that met your quality bar.
Two rules keep the number honest. Failed attempts, retries and abandoned runs stay in the numerator, because you paid for them. And the denominator counts only tasks that passed a check you would defend to a reviewer, such as an eval score, a human approval or a ticket that stayed closed. Track the number per workflow, not as one company-wide average, because an average hides the expensive workflows.
If you also need the return side of the equation, the guide to measuring enterprise AI cost savings covers baselines, payback period and attribution.
What to log on every request
To compute cost per task you need, for every model call: the model and pricing tier, input, cached and output tokens (with reasoning tokens broken out where the API reports them), tool and search calls, latency, a team or workflow tag, and the ID of the task the call belongs to.
Provider tools cover part of it. Anthropic's Usage and Cost Admin API returns historical usage and cost data that you can filter by API key, workspace, model and service tier, and OpenAI's Admin APIs cover spend limits and alerts, project administration and rate limits. Tying calls to tasks takes tracing in your own stack or an observability tool; see the LLM observability explainer for what to capture.
If you have no logs yet, you can still size the problem. Coworker's LLM cost calculator estimates a monthly bill across Claude, GPT, Gemini, DeepSeek and Kimi from the number of people using AI, tasks per person per day and typical task size.
Put a quality gate in front of every cost change
Every lever below trades something. Microsoft's Azure guide recommends running your eval set before deploying any prompt, model or routing change, and reverting if the score drops beyond a set tolerance, such as one point on a 100-point scale. It also recommends shipping no more than two cost changes at a time and measuring them for 7 to 14 days, because bundled changes make a regression hard to attribute. I would treat both as rules.
Coworker
See what your AI stack really costs
Compare model and platform costs, then run it all in one place.
Open the free LLM cost calculatorWhat are the most effective AI cost optimization levers?
These are the nine levers I would work through, starting at the request and ending with governance. OpenAI's own cost optimization guide lists the request-level basics in the same spirit: fewer requests, fewer input tokens, shorter outputs and smaller models, plus Batch and flex processing. The Tokenomics Foundation's Five-Layer Tokenomics Stack frames the split well: what a token costs is decided in the lower layers (silicon and capacity, which your providers own if you buy through APIs), and how many tokens you spend is decided in the upper layers (inference, model choice, routing and governance), which are yours.
1. Route each task to the cheapest model that passes your eval
Not every task needs the most capable model. Microsoft's Azure guide suggests a smaller, cheaper model for first-pass classification or extraction, escalating to a frontier model only on the requests where the small model is uncertain. The price gap in the token table above is 10x to 100x, so every task that stays on the small model counts.
The research backs this up. The FrugalGPT paper found that a learned cascade of models could match the performance of the best individual model, GPT-4 at the time, with up to 98% lower cost. LMSYS reported that its RouteLLM routers cut costs by over 85% on MT Bench, 45% on MMLU and 35% on GSM8K compared with using only GPT-4, while still reaching 95% of GPT-4's performance. Microsoft's Azure guide says a rule-based or learned router can send 60 to 80% of traffic to a cheaper model with no measurable quality drop.
Enterprises are acting on it. The State of Tokenomics survey found 86% of respondents evaluating or using a model router, and those using one were 4x more likely to be able to show value to the CFO. In its September 2026 AI Index, Ramp's lead economist reported hearing from businesses that impose company-wide defaults to reduce use of frontier models, because standard models are still highly performant and more cost effective.
To make it concrete, take 100,000 tasks a month at 4,000 input tokens and 1,000 output tokens each, priced at Anthropic's list rates:
| Setup (illustrative) | Monthly cost |
|---|---|
| Every task on Claude Opus 5.5 | $3,600 |
| Every task on Claude Haiku 4.5 | $900 |
| 70% on Haiku 4.5, 30% escalated to Opus 5.5 | $1,710 |
| Same split, with a 3,000-token shared prefix read from cache | $1,179 |
The 70/30 split is an assumption, not a benchmark, and cache writes are left out for simplicity. Your eval decides the real split. Revisit the routing table every quarter: Microsoft's guide notes that a model that was your only option six months ago might now cost five times more than a newer one that scores within one to two points on your evals. The explainers on LLM gateways and multi-model AI cover how routing is built, and a managed LLM gateway with smart routing is one way to buy it instead of building it.
2. Send less context: retrieve facts, not documents
The cheapest token is the one you never send. When an assistant or agent answers a question by pulling whole documents, threads and API responses into the prompt, the bill grows with every page. Past a point, more context can also lower answer quality, an effect covered in the context rot explainer.
The fix is to retrieve the smallest set of facts that answers the question. That can mean tighter chunking and reranking in a retrieval pipeline, summarizing tool results before they return to the model, or a memory layer that has already extracted and linked the facts, so an agent pulls a few relevant statements instead of re-reading the sources. The guides to context engineering and AI agent memory go deeper. Coworker's OM2 works this way, and its benchmark numbers are further down.
3. Trim tool definitions and tool results
If your agents connect to many tools through MCP or function calling, count the tokens those definitions add to every request. Anthropic's Tool Search Tool, which loads tool definitions on demand instead of all upfront, cut token usage by 85% in its example while keeping access to the full tool library. Presenting tools as code files the agent reads only when it needs them reduced one of its workflows from 150,000 tokens to 2,000, a 98.7% saving. Programmatic tool calling, which keeps intermediate results out of the model's context, cut average usage on complex research tasks from 43,588 to 27,297 tokens, a 37% reduction.
The practical steps are simple. Expose only the tools a given agent needs, keep tool descriptions short, and have tools return summaries or IDs instead of full payloads. The guide on how to choose MCP servers covers the scoping side.
4. Cache the stable part of every prompt
Prompt caching bills a fraction of the input rate for a prefix the provider has already processed. The rules differ by vendor.
| Prompt caching rule | Anthropic | OpenAI |
|---|---|---|
| How it turns on | Automatic caching or explicit breakpoints | On by default for supported models |
| Cache write | 1.25x the input rate (5-minute lifetime) or 2x (1-hour lifetime) | 1.25x the input rate on GPT-5.6 and later |
| Cache read | 0.1x the input rate on most models, 0.05x on Opus 5.5, 0.025x on Fable 5.1 and Mythos 5.1 | 0.1x on most GPT-5.6 and later models, 0.05x on GPT-6.1 Sol; discounts up to 95% |
| Minimum cacheable prefix | Varies by model | 1,024 tokens on GPT-5.6 and later |
| Lifetime | 5 minutes by default, refreshed on each use; 1 hour optional | At least 30 minutes after the last write or reuse on GPT-5.6 and later |
Sources: Anthropic prompt caching and OpenAI prompt caching, checked October 6, 2026.
OpenAI's documentation shows the payoff: writing a prefix once and reusing it nine times costs 2.15x its ordinary input cost, against 10x with no caching. The same arithmetic is why the Claude Code session above cost $0.55 instead of about $3.05.
Caching only works when the prefix is identical, so put stable content first (system prompt, tool definitions, reference material) and variable content last. Anthropic's documentation warns that a timestamp inside the cached block changes the prefix on every request, so you pay for a fresh cache write each time and never get a read. Track your hit rate; OpenAI provides a prompt caching dashboard for exactly that.
5. Batch, or use flex processing, for anything that can wait
Both major providers discount asynchronous work by half. Anthropic's Message Batches API charges 50% of standard prices on input and output, with most batches finishing in under an hour and results available within 24 hours, and caching discounts stack on top. OpenAI's Batch API also offers 50% lower costs with a 24-hour completion window and a separate pool of higher rate limits. OpenAI's flex processing, in beta with limited model availability, prices regular requests at Batch API rates in exchange for slower responses and occasional resource unavailability. Microsoft's guide notes the same 50% discount on the Azure OpenAI Batch API.
Good candidates include evaluations, data enrichment, nightly summaries, classification backfills, document processing and anything a person will read the next morning. One caveat from Anthropic: cache hits inside batches are best effort, typically between 30% and 98% depending on traffic patterns.
6. Cap output and reasoning tokens
Output is the expensive token, priced at 5x input on every model in the token table. Set a maximum output length on every call; OpenAI's reasoning guide points to the max_output_tokens parameter, which caps reasoning and visible output together. Lower the reasoning effort for routine work, since OpenAI describes lower effort as favoring speed and lower token usage, and Anthropic's docs suggest starting simple tasks near the 1,024-token minimum thinking budget and increasing it incrementally. Then ask for the format you need, such as a JSON object or a short list, instead of an essay you will trim.
Watch the failure mode. If a cap is too tight, a reasoning model can spend its budget thinking and return nothing visible, and you still pay for it. Size caps from logged usage, not guesses.
7. Put ceilings on agent loops and retries
Agents are where budgets get surprised, because a loop that does not converge keeps calling tools and re-reading context. Anthropic is direct about the trade-off in its multi-agent write-up: for economic viability, multi-agent systems need tasks where the value is high enough to pay for the extra tokens.
Set per-run ceilings on turns, tool calls and spend, and stop the run when one is hit. The Five-Layer Tokenomics Stack puts budgets, caps and circuit breakers at the routing and governance layer. Retry only the failed step, with backoff, instead of restarting the whole chain. And alert on drift: the FinOps Foundation's guidance is that if token consumption doubles without a clear reason, you should check for inefficient prompts or errors in application logic.
8. Govern seats: right tier, reclaim, consolidate
Seats are the easiest money to recover and the easiest to ignore. Review utilization monthly and reclaim seats that sat unused for a full billing cycle. Give premium seats, which cost 5x a standard seat on both Claude Team and ChatGPT Business, only to people whose usage justifies them. Consolidate overlapping assistants bought by different departments, and route expense-card AI purchases through procurement, since Zylo found ChatGPT is now the most expensed app.
Watch the hybrid plans. Where a seat includes metered usage, set per-user limits. Anthropic's Spend Limits API lets Claude Enterprise admins set a limit for each member, see where each member's limit is inherited from, and approve or deny requests for a higher one.
9. Allocate, budget and alert
None of the levers stick without ownership. The FinOps Foundation recommends a showback model, which shows each team its AI costs without immediately charging them as a chargeback model would, and pairing usage limits and throttling with anomaly detection. Tag every request with team, workflow and environment so the spend can be allocated at all.
Then use the controls your providers already ship. OpenAI's Admin APIs let you set an organization-wide monthly hard spend limit, after which affected requests return a 429 error, plus project spend alerts and per-project model allowlists. Anthropic's workspaces separate API keys, members and resource limits by project, team or environment under one bill. Microsoft's guide suggests budget alerts at 50%, 80% and 100%, routed to the same channel as your incident alerts.
Ownership is what makes the controls stick. In the State of Tokenomics survey, organizations with defined ownership of AI economics were 3.7x more likely to show value to the CFO, and none of the 12% with no owner could connect AI spend to a CFO outcome.
AI cost optimization levers compared
| Lever | What it changes | Best for | Effort | Impact reported by a source |
|---|---|---|---|---|
| 1. Model routing | Price per token | Mixed workloads with many routine tasks | Medium | Up to 98% lower cost at GPT-4 performance (FrugalGPT); 35% to over 85% lower, depending on benchmark (RouteLLM) |
| 2. Less context | Input tokens per task | Assistants and agents that read documents and threads | Medium to high | 66% to 89.1% lower token cost in Coworker's 100-task benchmark, depending on configuration |
| 3. Leaner tools | Input tokens per request | Agents with many MCP servers or functions | Low to medium | 85% fewer tokens with tool search; 37% with programmatic tool calling (Anthropic) |
| 4. Prompt caching | Price of repeated input | Long, stable system prompts and shared documents | Low | Cache reads at 10% of the input rate on most Claude models; up to 95% off cached input (OpenAI) |
| 5. Batch and flex | Price per token | Work that can wait from minutes to 24 hours | Low | 50% off (Anthropic, OpenAI, Azure OpenAI) |
| 6. Output and reasoning caps | Output tokens per task | Reasoning models and verbose responses | Low | No published figure; output costs 5x input |
| 7. Agent ceilings | Runaway loops and retries | Agents and multi-agent systems | Low to medium | No published figure; agents use about 4x chat tokens and multi-agent systems about 15x (Anthropic) |
| 8. Seat governance | Seat count and tier | Organizations with several AI subscriptions | Low | No published figure; 36% of SaaS licenses unused on average (Zylo) |
| 9. Allocation and budgets | Accountability | Every organization | Medium | No published figure; defined owners 3.7x more likely to show CFO value (State of Tokenomics) |
Effort ratings are my judgment. The impact figures come from different workloads and do not add up. The Coworker figures are statistically significant benchmarks comparing Coworker MCP vs. Claude Native Tooling; the 89.1% applies only to retrieval and context-heavy tasks with Coworker Learning, as the section below explains.
Which AI cost management tools cover which problem?
Tools split into categories by what they can see, and no single category covers seats, tokens and context at once. It also helps to know what buyers say they need: in the State of Tokenomics survey, only 4% wanted cheaper prices from model and token providers, while 23% asked for more transparency and granular data, 19% for standards such as FOCUS and 17% for attribution and tagging.
| Category | Examples | What it gives you | What it leaves to you | Pick it when |
|---|---|---|---|---|
| Provider consoles and admin APIs | Anthropic Console and Usage and Cost API, OpenAI Admin APIs | The provider's own usage and cost data, plus spend limits and alerts | A cross-provider view and seats bought elsewhere | One or two providers carry most of your API spend |
| FinOps platforms with AI support | CloudZero, Vantage, Harness, Flexera | AI provider spend allocated to teams and products alongside your cloud bill | Changing what each prompt sends | AI spend spans several vendors and a large cloud bill |
| LLM gateways | LiteLLM, OpenRouter | One endpoint in front of many models; LiteLLM tracks spend and sets budgets by virtual key, team or tag, and OpenRouter falls back to other providers when one goes down | Traffic that bypasses the gateway, and the size of each prompt | Engineering wants routing and per-team budgets in the request path |
| LLM observability | Langfuse | Usage and cost for every LLM call, filterable by user, feature or tag | Enforcement, and spend outside instrumented apps | You need to find which step of a workflow is expensive |
| SaaS management | Zylo, Flexera | Seat utilization, plus expensed or shadow AI apps | Token usage inside API workloads | Seat sprawl is the bigger problem |
| Context layer | Coworker (OM2, with model routing alongside) | Smaller, ranked context for each request | Cross-vendor spend reporting and seat governance | Agents and assistants read many sources per task |
Each vendor description comes from its own page, checked October 6, 2026. Vantage, for example, lists native integrations with OpenAI, Anthropic and Cursor with token visibility by developer, model and project, and Harness lists token and inference spend by agent, model and team across OpenAI, Anthropic, Bedrock and Vertex AI. The State of Tokenomics report found OpenRouter and LiteLLM had the highest adoption among routing tools, alongside many homegrown and cloud-native options. For a closer look at a developer router versus a full platform, see this OpenRouter comparison.
My honest take: if one provider carries most of your API spend, its console plus a gateway covers a lot before you buy anything. If AI is now a big share of a multi-cloud bill, a FinOps platform earns its cost. If seats are the problem, SaaS management is a better first purchase than any token tool. A context layer pays off only when context is a large share of the bill, so check your token breakdown first.
A 30-day AI cost optimization plan
This plan assumes you have API workloads and seat licenses but no dedicated AI FinOps practice yet.
Week 1: Get one number
- Pull 30 days of usage and cost from every provider console and cloud bill, and a seat roster for every AI tool, including expensed ones.
- Tag requests by team and workflow, and list your five most expensive workflows.
- Define cost per completed task for each of those five, including the quality check that counts as completed.
Week 2: Take the low-risk discounts
- Fix prompt structure so stable content comes first, turn on caching where it is off, and check the hit rate.
- Move evaluations, enrichment and other asynchronous jobs to a batch or flex tier.
- Set output caps and lower reasoning effort where your evals allow.
- Reclaim unused seats and move idle premium seats to standard.
Week 3: Route and trim on the top workflows
- Pick two levers per workflow, such as routing and context reduction, and ship them behind a feature flag.
- Measure for 7 to 14 days against your eval gate, then keep what holds and revert what does not.
- Remove tools that agents rarely call, and shorten the descriptions of the rest.
Week 4: Lock in guardrails and reporting
- Set spend limits and alerts at 50%, 80% and 100%, plus per-run ceilings for agents.
- Send each team a monthly showback report with cost per task for its workflows.
- Name one owner for AI spend overall and one per major workflow, and book a monthly review.
Mistakes that keep AI bills high
Chasing a lower token price instead of fewer tokens per task
A discount on a token you did not need is still waste. If an agent re-reads 100,000 tokens of context on every step, a 20% rate cut matters less than cutting the context in half. Buyers seem to know this: only 4% of State of Tokenomics respondents asked providers for cheaper prices.
Cutting model quality without an eval gate
A cheaper model that fails more often moves cost into retries and human rework, which your token report will not show. Measure cost per completed task before and after every change.
Budgeting seats and APIs separately
Some seat plans now carry metered usage, and the same piece of work can move between a seat-licensed assistant and an API-based agent. Budget them together, or one line will look healthy while the other absorbs the growth.
Setting caps with no owner
A cap with no owner either stops a working process or gets raised without a question. Every limit needs a person who can explain the spend behind it.
Where Coworker fits, and where it does not
Coworker works on the context and routing side of the bill. Its homepage says two pieces of infrastructure do the heavy lifting: "a portable context layer and an intelligent routing layer across open and closed models." The context layer is OM2, Coworker's organizational memory. It decomposes documents, messages and tickets from 50+ connected tools such as Salesforce, Slack, Jira and Google Drive into atomic facts, links them to entities such as people, companies, projects and deals, and scores them by proximity, recency, entity centrality and relationship strength. The graph narrows context before any model sees it. Access policies are inherited from the source tools, so people and agents only recall what they are allowed to see there. OM2 works through Coworker MCP in Claude, ChatGPT, Gemini, Perplexity, Cursor and custom agents, as well as in Coworker's own apps. The routing layer covers models from Anthropic, OpenAI and Google, plus open-weight models from Moonshot and Z.ai.
Coworker published a benchmark of this approach on its benchmarks page: 100 tasks across 7 categories (business operations, engineering, general business, marketing, people, product and sales), each repeated until statistically significant, on a Claude harness connected to Google Drive, Gmail, Google Calendar, HubSpot, GitHub, Jira, Slack, Notion and Stripe. The baseline was Claude with its own native connectors. The comparison was the same Claude working through the Coworker MCP. Cost is total input and output tokens per session.
| Configuration (Claude + Coworker MCP vs Claude with native connectors) | Token cost | Time to complete a task |
|---|---|---|
| OM2 alone, all 7 categories | 66% lower | 20% faster |
| OM2 with Coworker Learning, all 7 categories | 75.5% lower | 45.6% faster |
| OM2 with Coworker Learning, retrieval and context-heavy tasks (best case) | 89.1% lower (9x cheaper) | 64% faster |
| OM2 plus model routing | 98% lower (51x cheaper) | Not reported |
Statistically significant benchmarks comparing Coworker MCP vs. Claude Native Tooling. Coworker Learning is the part of OM2 that improves retrieval the more a team uses it. The 9x figure involves no model routing; the 51x figure is with OM2 plus model routing. Reviewers comparing answers blind, in randomized order, against a human-written reference preferred the Coworker answers 84.5% of the time. On roughly 5% of questions the native-connector baseline could not answer at all; those were excluded from the cost and speed comparisons and kept in the quality comparison.
Two caveats. These are Coworker's own numbers on a standard Claude harness and the sources listed above, and your mix of tasks will produce a different result. And a context layer does nothing for seat sprawl, GPU training costs or a workload made of short prompts with no retrieval; if that describes your bill, start with levers 1, 4, 5 and 8. On compliance, Coworker is SOC 2 Type II, GDPR and CASA Tier 2 compliant, runs US-hosted models and does not train on customer data. Alex Calder, Coworker's CEO, explains the thinking behind the product in introducing OM2 and changing the equation on enterprise AI spend.
Book a demo to see what a context layer and model routing do to your own AI bill. Coworker will also run its benchmark on your company's data and hand you a Context Impact Report, and it covers the token costs.
Frequently asked questions
What is AI cost optimization?
AI cost optimization is the practice of lowering what an organization pays for AI, including model API tokens, seat licenses and supporting infrastructure, without lowering the quality of the output. It is best measured as cost per completed task rather than cost per token. The main levers are model routing, sending less context, prompt caching, batch processing, output limits, agent loop ceilings, seat governance, and budgets with clear owners.
What is the difference between AI cost management and AI FinOps?
AI cost management is the visibility and governance work: knowing what AI costs, who spends it and whether it stays within budget across seats, APIs and cloud. AI FinOps, which the FinOps Foundation calls FinOps for AI, is the operating model behind it, bringing engineering, finance and business stakeholders together with showback or chargeback and unit metrics such as cost per inference. LLM cost optimization is the engineering work that changes how many tokens each request uses.
How do you reduce LLM costs without hurting quality?
Put an eval gate in front of every change, then start with levers that do not touch answer quality: prompt caching, batch processing for work that can wait, and trimming unused tool definitions. Next, route routine tasks to cheaper models and reduce the context each request carries, comparing eval scores before and after each change. Microsoft's Azure guidance suggests shipping no more than two cost changes at a time and measuring them for 7 to 14 days.
How much does prompt caching save?
On Anthropic's API, cache reads cost 10% of the normal input rate on most models and 5% on Claude Opus 5.5, while writing to the cache costs 1.25x the input rate for a 5-minute lifetime. OpenAI enables caching by default on supported models and discounts cached input by up to 95%. In OpenAI's own example, a prefix written once and reused nine times costs 2.15x its ordinary input cost instead of 10x.
When should you use a batch API?
Use it for any job that does not need an answer within minutes, such as evaluations, data enrichment, document processing, nightly summaries and classification backfills. Anthropic and OpenAI both charge 50% of standard prices for batch requests and return results within 24 hours, and Anthropic says most of its batches finish in under an hour. OpenAI's flex processing applies batch rates to regular requests that can tolerate slower responses.
Why are AI costs rising when token prices keep falling?
Volume is growing faster than prices are falling. Menlo Ventures predicted that net spending on generative AI would keep rising despite falling inference costs, driven by orders-of-magnitude growth in inference volume. Agents add to it: Anthropic found agents typically use about 4x the tokens of a chat interaction, and multi-agent systems about 15x.
How do you calculate AI cost per task?
Add up everything a task consumed: input, cached, output and reasoning tokens at their respective rates, plus tool and search fees, including failed attempts and retries. Divide by the number of tasks that passed your quality check. Track it per workflow instead of as one average, because an average hides the workflows that cost the most.
Are AI seat licenses still a fixed cost?
Less than they used to be. Claude Enterprise is priced as a seat plus usage at API rates, Claude Team and ChatGPT Business offer premium seats at five times the standard price with 5x the usage, and OpenAI lists credit-based and token-based pricing for ChatGPT Enterprise. Budget seats as partly variable and set per-user limits where the plan allows.
Related reading
- Enterprise AI pricing compared: 12 tools
- Enterprise AI cost savings: what to expect and how to measure it
- The Enterprise AI Price Index
- How much should enterprise AI cost per user?
- Changing the equation on enterprise AI spend
- What is an LLM gateway?
- What is a context window?
- What is context rot?
- What is LLM observability?
- LLM API cost calculator
- OM2 benchmarks
- Coworker pricing
Ready to get started?
Put Coworker to work inside your actual stack
Connect Salesforce, Slack, Jira, whatever you use, and run your first agent in minutes.