🧠 Introducing OM2: Your enterprise just started thinking.Read more
Benchmarks

100 tasks, measured. We'll run them on your company's data for free.

Across 7 categories, we measured what OM2 does to cost, speed, and quality against Claude's native connectors. The results and full methodology are below. We can run the same report on your company's data, and share the benchmarks with you for free.

Benchmark run

Run 3 of 3

100 tasks · 7 categories · measuring cost, speed and quality

Task

Baseline

Coworker

Won

Summarize Q3 pipeline

12.4s

24¢

4.4s

2.7¢

Find accounts at churn risk

22.8s

47¢

7.9s

5.2¢

Draft the renewal email

9.4s

18¢

3.6s

2.0¢

Who owns ACME-431?

no answer

2.1s

1.4¢

Reconcile last month's invoices

31.0s

62¢

10.6s

6.9¢

Which PRs touched billing?

15.6s

31¢

5.7s

3.4¢

Summarize the board deck

18.2s

39¢

6.4s

4.3¢

Top five support escalations

20.4s

42¢

7.2s

4.7¢

Status of the SOC 2 renewal

13.8s

27¢

5.0s

3.0¢

Who asked for SSO?

26.5s

54¢

9.3s

6.0¢

Open platform roles

11.2s

21¢

4.1s

2.3¢

100 of 100 complete

Scored blind, verified by a Coworker engineer

Time and token cost per task, with the blind pairwise winner. Baseline is Claude with its own native connectors. Coworker is the same Claude, through the Coworker MCP.

Results

Lower cost. Less waiting. Better answers.

Cost drops with every lever.

Cost is total input and output tokens per session. The first three bars are context optimization alone, no change to your model or workflow. The last adds Coworker's Optimized Routing on top.

Cost: tokens per session

2.9x cheaper

OM2 on Claude

vs. native connectors

66% cheaper

4.1x cheaper

+ Coworker LearningRoughly 40% of conversational prompts already invoke Coworker Learning successfully in production today, and for automation-style agents that figure is above 90%. Learning helps most where agents would otherwise rebuild the same queries every time, so the gains are largest on context-heavy workloads.

all 7 categories

75.5% cheaper

9.17x cheaper

+ Learning, retrieval

context-heavy tasks

89.1% cheaper

Coworker routes to the most efficient model

51x cheaper

+ Optimized Routing

stacked on top

98% cheaper

How we ran it

100 tasks across 7 categories, on a Claude harness with real connected tools. Each task was repeated until statistically significant, then replicated across every category: business operations, engineering, general business, marketing, people, product, sales. The connected sources were Google Drive, Gmail, Google Calendar, HubSpot, GitHub, Jira, Slack, Notion, Stripe.

Context Impact Report

Two things from you. We do the rest.

Our app runs the benchmark on your own systems, we cover the token costs, and you keep the numbers whichever way they land.

You

Kick off

Coworker

Connect · Build your graph · Run · Verify · Tune · Repeat

You

Read the numbers

Your part

Two steps

  • Sign an NDA and tell us which systems to cover.
  • Name one or two testers and have them install the Coworker benchmarking app.

Nothing to build, no scripts to write, and nothing for your engineers to maintain afterwards.

Our part

Everything else

  • Connect your sources and build OM2 on your company's data.
  • Run the full question set twice, once against your existing Claude setup and once through the Coworker MCP.
  • Score every answer blind and verify the results ourselves.
  • Tune your graph and run it again until the number settles.

We validate, adjust your graph and run it again. The last run is your number.

The ROI calculation was straightforward. We just needed to save two to two and a half hours per month per person to break even. We exceeded that in the first few weeks.

Marcin Safranow

Marcin Safranow

VP IT Operations, Huuuge Games

Huuuge Games

FAQ

Frequently asked questions

We ran 100 tasks, repeated until statistically significant, across 7 categories (business operations, engineering, general business, marketing, people, product, sales) on a Claude harness against 8 connected sources: Google Drive, Gmail, Google Calendar, HubSpot, GitHub, Jira, Slack, and Notion, plus Stripe. Cost is total input and output tokens per session, speed is end-to-end time to completion, and quality is a blind human pairwise preference against a human-authored reference answer, with order randomized.

Coworker Learning is the part of OM2 that improves retrieval the more your team uses it, instead of rebuilding the same queries from scratch each time. It is what takes the baseline 66% cost reduction to 75.5%, and up to 89.1% on retrieval-heavy tasks.

No. The 9x (9.17x, or 89.1% lower token cost) and 64% faster figures come from context optimization alone: no change to your model, no change to your workflow. Add Coworker's Optimized Routing across open and closed models on top of that, and total token cost drops 98%, or 51x cheaper overall.

On roughly 5% of questions, the native-connector baseline could not answer at all. Those were excluded from the cost and speed comparisons so the baseline was not unfairly deflated, but they were kept in the quality comparison, since failing to answer is itself a quality signal.

See what OM2 does for your team

The numbers above come from a standard Claude harness. We will run the same benchmark on your company's data and hand you a Context Impact Report. We cover the token costs.