Every model your engineers need, optimized under a ceiling you set
Engineering is where AI spend concentrates, and the usual fix is to cap access, which slows your best people down. Code lets engineers pick a model that fits the job, from near-free open models to frontier, keeps your company's context in the loop, and gives admins per-model caps and per-person visibility.
Coworker Code v1.0.2
Open Models gateway · Coworker credits
~/acme-payments
Your company's context is loaded. Change model any time with /model.
>/model
⎿Select a model. Price per 1M tokens, input / output
GLM 5.3 FlashZ.ai$0.15 / $0.50
Kimi K3Moonshot$3.00 / $15.00
Opus 5.5Anthropic$4.00 / $20.00
GPT 6 AstraOpenAI$10.00 / $50.00
⏺Switched to Kimi K3 for this session
>Tidy up the retry helper and add tests.
⏺Editedsrc/lib/retry.ts
⏺Createdsrc/lib/retry.test.ts
⏺Ranpnpm test⎿ 48 passed
· Done in 22s · 0.4% of your monthly credits used
Pick the model that fits the job.
Where the AI bill goes
Across the CIO conversations we have had
What OM2 + routing does to that bill
*Statistically significant benchmarks comparing Coworker MCP vs. Claude Native Tooling
What your CIO gets to decide
Lower spend without slowing anyone down
Engineers can easily move routine work to open models that cost a fraction of frontier ones, so your AI bill comes down without anyone losing access to the models they need.
Spend you can see and limit
Set usage limits by model and by user, across Code and our other native apps, with exemptions for the engineers you cannot afford to slow down.
Nothing to migrate
Engineers keep Claude Code and the skills and history they already have with it. What changes is what runs underneath it, so there is no rollout project and no retraining.
Cheaper models without a procurement cycle
Open models are run on secure US-hosted providers that are SOC 2, GDPR and ISO certified. Code uses the same usage-based credits as the rest of Coworker, so there are no new commercial conversations to have.
Why engineers love
Coworker 😍
A routing layer doesn't know your company
Every tool in this category can send a request to a cheaper model. The question is what the model knows when it gets there.
A routing layer sees
No other context available
So “add retries” is all it ever knows. A cheaper way to buy tokens for one tool.
Code sees
- #eng-paymentsSlack thread
- Payments v2 kickoffCall transcript
- Checkout resilienceDesign doc
- PR #2211Review
Idempotent retries on the Stripe webhook only.
Same ticket. The scope lives in the thread, the call and the doc, so that is where it goes looking.
From install to a bill you can see
- Step 1
Install Code
One command on top of Claude Code and npm, which adds coworker to the terminal.
- Step 2
Log in and pick your model
Paste the API key from the Coworker app, then choose GPT, Gemini, Kimi, GLM or Claude, or set presets for the fast and smart slots.
- Step 3
Start coding
Type coworker and work the way you already do in Claude Code. Every session starts with your company's context already loaded, which plain Claude Code cannot do, because Code and the Coworker MCP were built to work together.
Questions
Frequently asked questions
You keep it. Bring your own Anthropic key and your Claude spend stays exactly where it is, while everything else goes through us.
Yes, at any point. Swap models mid-task, or switch between Code and Claude Code and pick the same session back up, so the work carries across. Reaching a cap on one model leaves every other model available, so the work keeps moving.
Anywhere you have a terminal. It is a terminal app, so it runs in whatever terminal your engineers already use, and in the integrated terminal inside editors like VS Code, JetBrains and Cursor.
Yes. Each request goes to the model the engineer selected, and the tool picker translates tool calls into the format that model was trained on, so Kimi and GLM behave correctly inside a coding agent even though they were not built for one. Skills port across as well.
Open models run on US-hosted inference providers that are SOC 2, GDPR and ISO certified, currently Fireworks and Baseten. We host all open models in the US to meet residency requirements, and Coworker is SOC 2 Type II certified.
In about 30 seconds. Getting GLM or Kimi approved usually means days of vendor review and a new contract. Here it is one install, on US-hosted providers we already run under SOC 2, GDPR and ISO.
No. Bedrock keeps serving Claude exactly as it does today, and Code runs alongside it for everything else. We do not touch your Claude Code install.
No. Caps and allow lists are enforced at the gateway rather than on trust, and usage lands per person and per model in the admin view.
On the same usage-based credits as the rest of Coworker, so it draws on the balance you already have rather than a separate line item.
Compare
See how Code compares
Coworker vs Cursor
Cursor's context stops at the repo and its model list is someone else's to choose. Coworker brings your company's context and every major lab.
Read comparison →Coworker vs Claude Code
Claude Code sells one lab's models. Code gives your engineers GPT, Gemini, Kimi, GLM and Claude, under caps you set.
Read comparison →Coworker vs Microsoft Copilot
Copilot knows your code. Code knows your meetings, tickets and docs too, and shows the spend per person.
Read comparison →