On this page
Put Coworker to work on your stack.
Connect Salesforce, Slack, Jira and run your first agent in minutes.
Enterprise AI
What Is a Context Layer? Context Layer vs Semantic Layer, and How to Build or Buy One
Coworker AI explains the context layer: what it is, how it differs from a semantic layer, RAG and knowledge graphs, and how to build, buy or test one.
A context layer is the part of an AI stack that sits between a company's data and tools and the AI models and agents that use them, and gives each request current, permission-aware context that is relevant to the task. It decides what a model actually sees: which facts, from which systems, for which user, and whether those facts are still true. A semantic layer tells AI what a metric means; a context layer covers the rest of what an agent needs before it answers or acts.
Without one, an agent rebuilds its picture of your company on every task through whatever tool calls it happens to make. It has no reliable way to know that a decision from Tuesday's call was reversed in Slack on Thursday, or that the person asking should not see the deal size. A context layer does that work ahead of time and keeps it current for every assistant and agent the company uses.
The term is young, and vendors use it differently. DataHub calls it a newer and less settled concept than the semantic layer. Data platforms such as DataHub and Atlan build it on governed metadata and business definitions, Redis uses the term for agent memory and session state, and enterprise search and organizational memory products apply it to knowledge spread across SaaS tools. This guide covers where they agree, how a context layer differs from a semantic layer, RAG and a knowledge graph, and the buyer problems I found thinnest in the pages that rank for this term today: permissions on combined facts, facts that change, conversations as a source, delivery over MCP, build vs buy, and testing on your own data.
I read every source cited here on October 6, 2026. I work at Coworker, which sells one kind of context layer, so the Coworker material sits in its own section near the end, together with the cases where other tools fit better.
What is a context layer?
A context layer is infrastructure, not a model or a prompt. It connects to the systems where a company's knowledge lives, organizes what it finds into facts and entities, keeps that picture current, applies each user's permissions and assembles the right slice for each AI request. You will also see it called an AI context layer. An enterprise context layer is the same idea built for a whole company rather than one agent or one team, and that is the version this guide focuses on.
DataWalk's shorthand is useful: data is the record, context is the meaning. A payment is a row in a table until something connects it to the customer, the account history and the activity that explains it.
What does a context layer do?
The vendor definitions differ in emphasis, but between them they describe six jobs:
- Connect. Read from the systems where work happens: warehouses, CRMs, ticketing tools, documents, chat, email and meeting notes.
- Resolve entities. Work out that "Acme Inc" in the CRM, "ACME" in a support ticket and "the Acme renewal" in Slack are one customer, and that a Slack handle and an email address belong to one person.
- Keep facts current. Pick up changes as they happen, and stop serving a fact once a later decision replaces it.
- Enforce permissions. Show each person, and each agent acting for them, only what that person can open in the source system.
- Rank and assemble. Pick the few facts that matter for this request and fit them into a token budget.
- Deliver. Hand the result to models and agents through an API or the Model Context Protocol (MCP), so the same context reaches every AI tool.
Where does the context layer sit in the AI stack?
Between your systems and your models. One way to draw the stack is five layers:
| Layer | What it does | Examples |
|---|---|---|
| Systems of record and work | Where data and decisions originate | Salesforce, Snowflake, Jira, Slack, Gmail, Zoom, Google Drive |
| Context layer | Connects, resolves entities, keeps facts current, applies permissions, ranks | Data context platforms, agent memory runtimes, organizational memory |
| Delivery | Moves context to AI tools | MCP servers, APIs, SQL |
| Models and routing | Generates the answer, and decides which model handles which request | Claude, GPT, Gemini, open-weight models, an LLM gateway |
| Assistants and agents | Where people and automations ask for work | Claude, ChatGPT, Cursor, custom agents |
DataWalk, summarizing a March 2026 Gartner report, says Gartner's framework places the context layer as its own tier between the information layer and the intelligence layer. Routing and context are separate jobs: routing decides which model answers, and the context layer decides what that model sees.
Why do AI agents need a context layer?
Three of the vendors that rank for this term make the same argument: the weak point is usually the input, not the model. Snowflake's Josh Klahr and Rajhans Samdani write in their agent context layer post that the bottleneck is going to be the right context. Redis says most agents fail in production because of the inputs they are handed, and names five failure modes: context poisoning, distraction, confusion, clash and rot. DataWalk says many enterprise AI failures come from missing or fragmented context rather than weak models.
Two pieces of public research explain why giving the model everything does not work either. Chroma tested 18 models in July 2025 and found that performance grows increasingly unreliable as input length grows. Anthropic's engineering team describes context as a finite resource with diminishing marginal returns, and good context engineering as finding the smallest set of high-signal tokens for the task. A context layer is the infrastructure that produces that small set at company scale. For the model side of this, see our explainers on the context window, context rot and contextual understanding.
Context layer vs semantic layer: what is the difference?
A semantic layer defines what a metric means and how it is calculated, so every dashboard and query returns the same number. A context layer covers what an AI system needs beyond that before it acts: who may see a fact, where it came from, whether it is still current, and what people decided about it in documents and conversations. Ataccama puts it in one line: the semantic layer says what data means, and the context layer says how that meaning applies.
| Semantic layer | Context layer | |
|---|---|---|
| Question it answers | What does this metric mean, and how is it calculated? | What does this agent need for this request, for this user, and is it still true? |
| Built for | Analysts, BI tools, text-to-SQL | AI assistants and agents working across tools |
| Main inputs | Warehouse tables, joins, metric definitions | Warehouse data plus documents, tickets, CRM records, chat, email and meeting notes |
| Output | One consistent number per metric | A ranked set of facts and entities, with sources, sized for a prompt |
| Permissions | Controls who can query which metrics and the data behind them | Per user and per source, including facts that combine sources |
| Changes when | Someone edits the metric logic | Any source changes, or a later decision replaces an earlier one |
| Conversations | Out of scope | In scope, since decisions get made in Slack threads, calls and email |
| When it fails | Two dashboards disagree | An agent cites an outdated or restricted fact with confidence |
| Typical tools | dbt Semantic Layer, LookML, Cube, AtScale, Snowflake semantic views | Data context platforms, agent memory runtimes, organizational memory |
What semantic layers do well
Semantic layers fixed a real and expensive problem: the same question returning different numbers depending on which tool you asked. DataHub cites GigaOm classifying semantic layers as a mature market. dbt's documentation describes the core benefit: define a metric once in the modeling layer, and when the definition changes it is refreshed everywhere it is used. Semantic layers already reach AI tools, too. dbt says Claude and ChatGPT can query its Semantic Layer through the dbt MCP server, so answers use governed metrics instead of guesses at raw tables.
If your agents mostly answer questions like "what was Q3 revenue in EMEA?", fix the semantic layer first. DataHub uses that exact question as its example of what a semantic layer answers well.
What a context layer adds
DataHub's second example is the one that needs a context layer: an agent asked to explain why EMEA revenue dropped and to recommend next steps. DataHub lists what that takes beyond metric definitions: lineage to trace the data, freshness signals, governance policies about what can be shared with whom, and business rules about acceptable recommendations. It also needs the conversations, if the reason lives in them: the Slack thread where sales flagged a delayed deal, or the call where a customer asked for a discount. Neither is a metric, and a semantic layer does not model them.
Do you need both?
If you run BI at any real scale, yes. The vendors disagree on how the two stack. DataHub argues they are interdependent rather than built one on top of the other, Redis draws the context layer above both RAG and the semantic layer, and Ataccama adds a data trust layer underneath. The order matters less than the boundary: keep metric definitions in the semantic layer your data team already governs, and have the context layer read from it instead of defining revenue a second time.
Atlan co-founder Prukalpa Sankar names the failure to avoid in Atlan's field guide to the enterprise context layer: separate layers for support agents, data analysis, memory and process mining become "context islands" that cannot share what they learn.
Context layer vs RAG, vector databases, knowledge graphs and agent memory
These terms often get compared to a context layer. Most of them are parts of a context layer rather than alternatives to it.
| Term | What it is | What it leaves out | Relationship to a context layer |
|---|---|---|---|
| RAG | A technique: retrieve passages and add them to the prompt | State across tasks, entity resolution, deciding what is outdated | One retrieval method a context layer can use |
| Vector database | Storage and similarity search over embeddings | Complete lists, counts and negation; entities; freshness rules | A storage component |
| Knowledge graph | Entities and typed relationships | Per-request ranking, permissions and delivery, unless added on top | Often the core data structure |
| Agent memory | What one agent remembers across turns and sessions | Facts the rest of the company should share | Working memory stays with the agent; shared knowledge belongs in the layer |
| Data catalog | An inventory of data assets, built for people | Conversations, documents, per-request assembly | A source of metadata |
| MCP | A protocol connecting AI apps to tools and data | Storage, ranking, permission logic | The delivery route |
| Context engineering | The practice of deciding what a model sees | The infrastructure itself | The discipline; the context layer is what it runs on |
Is RAG a context layer?
No. RAG is a retrieval technique: find passages similar to the question and put them in the prompt. Redis and DataWalk draw the same line. RAG retrieves for a single call, while a context layer also decides whether retrieved information is still valid, carries state across tasks and decides what to forget. A context layer often uses RAG inside it. Our RAG vs MCP explainer covers how retrieval and connection protocols fit together.
Is a knowledge graph a context layer?
Not on its own. A graph stores entities and the relationships between them, which is what makes multi-hop questions answerable, like "which customers are affected by an outage in the service Dana owns?" To act as a context layer it also needs connectors that keep it current, permissions inherited from each source, ranking for a token budget and a delivery route to AI tools. DataHub's FAQ says the same: a knowledge graph is one component that can contribute to a context layer. Our guide to the enterprise knowledge graph covers structure and limits.
A related term, "context graph," spread quickly after Foundation Capital's December 2025 essay by Jaya Gupta and Ashu Garg. Their version is a graph that also records decision traces: what was decided, under which policy, with whose approval, and which precedent it followed.
Is agent memory the same thing?
They overlap, and they are not the same. Agent memory is what one agent keeps across turns and sessions. Atlan's essay splits memory four ways. It places working memory and episodic memory (the record of what happened) with the agent, and semantic memory (durable facts) and procedural memory (how work gets done) in the context layer, because those are what other agents and people should share. For the agent side, see AI agent memory and our Mem0 vs Zep vs Letta comparison.
How does context engineering relate?
Context engineering is the practice, and the context layer is the infrastructure it runs on. Ataccama frames it the same way: if the context layer is the thing, context engineering is the work of building and maintaining it. Our context engineering guide covers the practice: what goes in the window, in what order and how much of it.
Who sells a context layer? Three kinds on the market
The products I reviewed for this guide fall into three groups, depending on what they were built from.
| Type | Examples | Built mainly from | Strongest at | Better fit when |
|---|---|---|---|---|
| Data context layer | DataHub, Atlan, Snowflake, Ataccama | Warehouse tables, metadata, lineage, glossaries, documentation | Governed metrics, lineage, data quality, analytics agents | Your agents mostly answer questions from structured data |
| Agent runtime context | Redis Iris, MindStudio | Session state, agent memory, caches, operational databases | Fast memory and state for agents you build | You are building your own agent product |
| Work context (organizational memory) | Glean, Coworker | Documents, tickets, CRM records, chat, email and meeting notes across SaaS tools | Knowledge about people, customers and decisions across tools | Your assistants and agents work across many SaaS tools and conversations |
Data context layers
DataHub, Atlan, Snowflake and Ataccama start from the data estate. DataHub describes a context graph that connects metadata from 100+ data sources, including Snowflake, Databricks, dbt, Looker, Notion and Confluence, kept current by an event-driven architecture and exposed to AI through MCP servers and connectors for Claude, Cursor and Snowflake Cortex. Snowflake's design is layered: an analytic semantic model, a relationship and identity layer, operational playbooks, provenance, and event and decision memory. Snowflake reports that in an internal experiment, adding a plain-text data ontology (join keys, table grains, fan-out hints) improved final answer accuracy by 20%, cut tool calls by about 39% and cut latency by about 20% against a best-practices baseline. Ataccama argues the most important layer sits underneath: data quality, entity resolution and lineage, which it calls a trust layer.
These are the better choice when your agents mostly answer questions over warehouse data: text-to-SQL, governed metrics, lineage. Conversations are where they differ most. DataHub mentions Slack as a place where business rules end up scattered, while Atlan's essay lists Slack, email and tickets among the systems a context layer should mine.
Agent runtime context layers
Redis defines the context layer as the part of the stack that decides what an agent knows at any moment, and sells Redis Iris to run it: session memory with time-based expiry, long-term memory with embeddings, a semantic cache, change data capture from operational databases such as Postgres and MySQL, and a Context Retriever that generates MCP tools from a semantic model, with row-level access controls enforced on the server. Redis reports up to 15x faster cache hits in its benchmarks and up to 73% lower LLM inference costs from the semantic cache. MindStudio bundles memory, retrieval, state management and permission scoping into its agent builder and lists 1,000+ integrations. These fit teams building their own agents who care about latency on every step. You still decide which organizational knowledge flows into them, and wire it up.
Work context layers
Glean and Coworker start from the SaaS tools where knowledge work happens, the approach behind organizational memory products. Glean's connectors page lists 275+ native and MCP-based connectors, says all data permissions are inherited and strictly enforced, and says permission changes in a source show up in results immediately. Its Enterprise Context page says Glean MCP delivers that context into Claude Code, Codex, Gemini, Cursor and Copilot. Glean's engineers also explain entity resolution well: their post on building a context graph describes how the knowledge graph learns that "ACME Inc" in the CRM and "ACME" in support tickets are the same customer.
Our Coworker vs Glean comparison covers how the two differ, and Glean Agent Builder covers Glean's agent side. Coworker's own approach is in the worked example below.
Coworker
Put Coworker to work on your actual stack
Connect Salesforce, Slack, Jira and run your first agent in minutes.
Permissions: who is allowed to see a combined fact?
This is the problem I would test first. None of the seven context-layer pages I read for this guide explains what happens to a fact built from sources with different access rules. Six of them mention access control in some form: most as an item on a list, MindStudio as a goal (assemble only the context the current user is entitled to see), and Redis as row-level access for database rows.
Inheriting permissions from one source is the easy case: if you cannot open a Google Doc, the assistant should not quote it to you. The hard case is a fact built from several sources. Take an account summary that combines:
- a Jira ticket about a bug, visible to the whole company
- a private Slack channel where the account team discusses churn risk
- a Salesforce opportunity visible only to sales
A support engineer asks, "what's going on with Acme?" They should get the bug and nothing about churn risk or deal size. A sales VP asking the same question should get all three. The same question from two people should produce two different answers. A layer that writes the combined summary once, and then checks only who may read the summary, has already leaked.
How should a context layer handle permissions on derived facts?
There are two common ways to apply permissions, and I would want a vendor to use both:
- Check at query time. When someone asks, filter candidate facts by what that person can open in each source right now. This is accurate at the moment of the question, at some cost in latency.
- Copy permissions at sync time. Store each source's access rules with the data and sync changes. This is fast, and only as current as the last permission sync.
For a fact derived from several sources, the safe rule is the intersection: show it only to people who could open every source it came from, or rebuild it per user from the sources that user can see. Then ask how quickly a permission change in the source reaches answers (someone removed from a private channel, an employee offboarded), and whether every answer keeps provenance so it can show which sources it used.
What changes when context is delivered through MCP?
Delivery adds a second identity question. The MCP specification says the protocol cannot enforce its security principles at the protocol level, so access control is the implementer's job, and its authorization section makes authorization optional. When an MCP server does use authorization over HTTP, the spec builds on OAuth 2.1, requires the server to check that each access token was issued for it, and forbids it from accepting or passing along any other tokens.
For a context layer the practical question is simpler: does each person connect as themselves, so the server knows whose permissions apply, or does everyone share one service account? A shared account flattens every permission you carefully inherited. Our guides to MCP security and enterprise MCP go deeper.
Freshness: what happens when a decision changes?
Every context layer is meant to stay current, and two different properties hide inside that word. They need separate tests.
Sync freshness vs fact freshness
Sync freshness is how quickly a change in a source shows up in the layer: a new Slack message, an edited doc, a closed ticket. Vendors describe several ways to get it: change data capture from databases (Redis), event-driven metadata updates (DataHub) and real-time sync that includes permission changes (Glean).
Fact freshness is whether the layer knows that a newer statement replaced an older one. That is harder, because the two statements can live in different tools, and neither says "this supersedes that."
Here is an example. On Monday's pricing call, the team agrees to cap the Acme renewal discount at 15%. On Wednesday, the CFO replies in a Slack thread: hold it at 10%. On Thursday, an account executive asks an assistant what discount is approved for Acme. The right answer is 10%, citing the Slack thread. A layer that ranks only by similarity to the question may return the meeting notes instead, because they discuss the Acme discount at length. Both sources synced on time, and the answer is still wrong.
What does good supersession look like?
The approaches in published material fall into four groups:
- Validity periods and history. DataWalk describes a context layer that keeps historical state, changing relationships and validity periods, so a new fact can replace an old one without erasing what was true before.
- Recency and confidence in ranking. Redis says a context layer should rank by recency and confidence so that contradictory fragments do not land in the prompt with equal weight.
- Versioned definitions with provenance. Snowflake's churn example returns the metric version (v2.4) along with what changed since the previous report.
- Change propagation. Atlan's essay argues that when a core definition changes, the update should trace through everything that depends on it, which it calls the blast radius of a change.
Ask a vendor to run the Acme example on your data: record a decision, reverse it in a different tool, then ask. See which answer comes back and whether it cites the newer source.
Why conversations belong in a context layer
Decisions get made in conversation, and conversations are among the hardest sources for a context layer to read well.
Foundation Capital's essay makes the case directly. Gupta and Garg argue that the exceptions, overrides and precedents that run a company live in Slack threads, deal desk conversations, escalation calls and people's heads. Their example is a VP approving a discount on a Zoom call or in a Slack DM, while the opportunity record shows only the final price. Glean's engineers make a similar point: the outcome of a decision may land in a system of record, but the work behind it happens in meetings, chat, email and documents.
If your agents need to know why something happened, and not only what the numbers say, the conversations are where much of that record lives.
Why are conversations harder than tables?
- They contradict each other, and the latest message is not always the final word.
- They refer to things loosely ("the Acme thing", "Dana's project"), which makes entity resolution harder.
- Their permissions are fine-grained: private channels, DMs, meeting invites, email recipients.
- They mix decisions with chatter, so a layer has to tell a commitment from a passing idea.
Which conversation sources should you connect first?
Start where decisions are recorded: chat (Slack, Microsoft Teams), email (Gmail, Outlook), and meeting notes and transcripts (Zoom, Gong, Granola, Fireflies.ai, Fellow). Then add the systems those conversations refer to: CRM (Salesforce, HubSpot), tickets (Jira, Linear, Zendesk, Intercom) and documents (Google Drive, Notion, Confluence, SharePoint). A layer that reads only the second group can tell you a deal closed at 10% off. It needs the first group to tell you who approved it and why.
How does a context layer reach Claude, ChatGPT and Cursor?
Through the Model Context Protocol or an API. The Model Context Protocol is an open-source standard for connecting AI applications to external systems, and its site lists Claude, ChatGPT, Visual Studio Code and Cursor among the clients that support it. The current specification, version 2026-07-28, defines three things a server can offer: resources (context and data), prompts (templated messages and workflows) and tools (functions the model can call). For a context layer, MCP is the delivery route: build the layer once, and every MCP client can query it.
Client support, from each vendor's documentation on October 6, 2026:
| AI tool | MCP support |
|---|---|
| Claude | Custom connectors over remote MCP on Free, Pro, Max, Team and Enterprise plans; Free is limited to one custom connector |
| Claude Code | Connects to remote MCP servers over HTTP; tool search defers tool definitions until they are needed, to keep context usage low |
| ChatGPT | Full MCP support, including write actions, rolling out in beta to Business, Enterprise and Edu plans through developer mode |
| Cursor | MCP servers over stdio, SSE or Streamable HTTP, with OAuth for the remote transports |
Can one context layer serve every AI tool?
That is the main reason to deliver context over MCP. If engineers work in Cursor and Claude Code while sales works in ChatGPT, a context layer behind one MCP server gives both groups the same governed facts, instead of a separate partial index inside each tool. Several context-layer vendors now ship MCP servers: DataHub exposes context through them, dbt offers one for its semantic layer, Glean MCP feeds Claude Code, Cursor and others, and Ataccama's MCP server exposes governed data with its trust signals. Our explainers on what MCP is, Claude MCP and ChatGPT MCP cover setup.
What does MCP not do for you?
MCP moves context. It does not decide what is true or who may see it, and the specification says as much about security. Four jobs stay with the context layer:
- Identity. Per-user authentication, so the server applies the right person's permissions.
- Freshness. An MCP server returns whatever the layer behind it holds, current or not.
- Selection. Tool definitions and tool results both take up the model's context window, which is why Claude Code defers loading tool definitions. A layer that returns a ranked handful of facts usually costs fewer tokens than one that returns raw documents for the model to read.
- Consistency. If each AI tool connects to a different set of servers, each gets a different picture of the company.
For governing many servers at once, see our guide to MCP gateways.
Build vs buy: what does it take to build a context layer?
Build if the scope is narrow (one or two sources, one use case) and a platform team will own it for years. Buy if you need many SaaS sources, conversations and per-user permissions across several AI tools. The components are the same either way, and much of the work is in keeping them running after launch.
| Component | What building it involves | Where the ongoing work is |
|---|---|---|
| Connectors | OAuth per source, pagination, rate limits, webhooks; reading content, metadata and activity, not only documents | Vendor API changes, new sources, quirks in each app |
| Entity resolution | Matching accounts, people and projects across tools: "ACME Inc" vs "ACME", a CRM account ID vs a support org ID | New aliases, renamed accounts, reorganizations |
| Permissions | Mirroring each source's access rules, groups, private channels and DMs; a rule for derived facts | Permission changes, offboarding, audits |
| Freshness | Change capture or polling per source; marking superseded facts | Sync lag, contradictions between sources |
| Ranking and assembly | Scoring relevance, recency and authority; fitting a token budget | Tuning as usage grows; long inputs degrading answers |
| Delivery | An MCP server with OAuth 2.1, plus APIs | Differences between Claude, ChatGPT, Cursor and other clients |
| Evaluation | Golden question sets, permission-leak tests, stale-fact tests | Re-running them after every change |
| Security and compliance | Encryption, audit logs, retention rules, certifications | Annual audits, customer security reviews |
MindStudio, which sells a platform for this, estimates three to six months of infrastructure work for most teams before anything useful runs on top. That is a vendor's estimate rather than a measurement, but the table above shows where the time goes.
When does building make sense?
- Your context is mostly in a warehouse you already govern. Start with your semantic layer or your data platform's context features rather than a new system.
- One agent needs one or two sources, and the permissions are simple.
- Context is part of what you sell. If your product is an agent, its context layer is product work, and runtime components like Redis are built for that job.
- You have a team that will own connectors, permissions and evaluation for years, not one quarter.
When does buying make sense?
- Your agents need many SaaS sources, including chat, email and meetings.
- Permissions vary widely by source and by person.
- You want the same context in several AI tools, not one.
- Your platform team is already stretched, and the business wants results this quarter.
What should you ask a context layer vendor?
- Which sources do you read, and do you ingest content, metadata and activity, or only documents?
- How do you handle a fact that combines sources with different permissions?
- How long does a permission change in a source take to reach answers?
- How do you detect that a newer statement replaced an older one, and can I see the history?
- Does every answer cite its sources?
- Which AI tools can use the layer, and does each user authenticate as themselves?
- Can I export my data, and anything derived from it, if I leave?
- Will you run an evaluation on my data, with my questions, scored blind?
For a broader procurement list, see the enterprise AI buyer's checklist.
How do you test a context layer on your own data?
Vendor benchmarks, Coworker's included, tell you how a system did on someone else's questions. Before you commit, run these four tests on yours.
Build a golden question set
Pull 50 to 100 questions from real work in the last quarter: Slack threads where someone asked if anyone knew something, escalations, deal reviews, onboarding questions. Write down the answer your team agrees on and where it lives. Include structural questions, because they separate a context layer from search: counting ("how many enterprise renewals slipped last quarter?"), lists ("which customers reported the SSO bug?"), negation ("which open deals have no executive sponsor?") and multi-hop ("who owns the services affected by last week's outage?"). Similarity search tends to return a few plausible examples where you needed a complete list.
Run a permission-leak test
Set up two test users with different access, such as a support engineer and a sales lead. Ask both the same 20 questions that touch restricted sources. A pass means neither answer contains or cites anything that user cannot open in the source tool. Then remove one user from a private channel and run it again.
Run a stale-fact test
Record a decision in one tool and reverse it a day later in another, as in the Acme example. Then ask. A pass means the newer decision wins and the answer cites it.
Measure cost, speed and blind preference
Run the same tasks through your current setup and through the candidate. Measure tokens per task and time to a finished answer. Have reviewers compare answers blind, in random order, without knowing which system produced which. Leave questions the baseline cannot answer at all out of the cost and speed comparison, but count them in quality, since failing to answer is a quality result.
| Test | What a pass looks like |
|---|---|
| Golden questions | Correct, cited answers on the large majority, with failures you can explain |
| Structural questions | Complete lists and correct counts, not a few similar-looking examples |
| Permission leak | No answer contains or cites a source that user cannot open |
| Stale fact | The newer decision wins, with the newer source cited |
| Cost and speed | Fewer tokens and less time per task than your current setup, on the same tasks |
| Blind preference | Your reviewers prefer the new answers without knowing which system wrote them |
If you would rather not build the harness yourself, Coworker will run one for you. Its free Context Impact Report runs the full benchmark question set on your company's data, once against your current setup and once through Coworker, and scores every answer blind. You sign an NDA and name one or two testers; Coworker covers the token costs, and you keep the numbers whichever way they land.
Worked example: how Coworker's OM2 works as a context layer
Coworker is my employer, so read this as the vendor's view and test it like any other. Coworker's homepage says two pieces of infrastructure do the heavy lifting: "a portable context layer and an intelligent routing layer across open and closed models." The context layer is OM2, Coworker's organizational memory. In the OM2 launch post, CEO Alex Calder wrote: "You can think of OM2 as a context graph." Here is how it maps to the requirements in this guide, using only what Coworker states publicly:
| Requirement | What Coworker says publicly |
|---|---|
| Sources | 50+ connectors, including Slack, Gmail, Zoom, Gong, Salesforce, HubSpot, Jira, GitHub, Google Drive and Notion |
| Structure | Documents, messages and tickets are decomposed into atomic facts, enriched with entities and relationships (organizational memory) |
| Ranking | Graph algorithms score context by proximity, recency, entity centrality and relationship strength |
| Freshness | OM2 continuously takes in new information and invalidates outdated information |
| Permissions | Access policies inherited from the underlying sources travel with every fact and connection, and source permissions are enforced on traversal |
| Provenance | Provenance chains back to the underlying sources, so every answer traces to its source |
| Structured questions | Enumeration, negation, counting, structural ranking and multi-hop traversal |
| Delivery | MCP or API, in Claude, ChatGPT, Gemini, Perplexity and custom agents; the Coworker MCP page also lists Claude Code, Cursor and Windsurf |
| Security | SOC 2 Type II, GDPR and CASA Tier 2; customer data is never used to train models; US-hosted models (platform) |
Coworker also published a benchmark of 100 tasks across seven business categories, run on a Claude harness with real connected tools. Claude using the Coworker MCP with OM2 (including Coworker Learning) had 9x lower token cost and finished tasks 64% faster than the same Claude with its own native connectors, on retrieval-heavy task categories, which is the best case. Across all seven categories the gains were smaller: 75.5% lower token cost and 45.6% faster. In blind side-by-side review, reviewers preferred the OM2 answers 84.5% of the time. Adding Coworker's model routing on top brings the cost figure to 51x cheaper with OM2 plus model routing. The method, including the roughly 5% of questions the baseline could not answer, is on the benchmarks page.
Source: Statistically significant benchmarks comparing Coworker MCP vs. Claude Native Tooling.
Where Coworker fits less well
If your agents mostly answer metric questions over a warehouse, a semantic layer or a data context platform such as DataHub or Atlan is the more direct fix. If you are building your own agent product and need in-memory session state on every step, runtime infrastructure like Redis is closer to that job. If you need the widest connector catalog, Glean lists 275+ connectors to Coworker's 50+. And Coworker's public pages describe what OM2 does, not every mechanism behind it, so run the permission-leak and stale-fact tests on it as you would on any other vendor.
To get numbers on your own systems, start with the free Context Impact Report. Book a demo to walk through Coworker on your stack.
If you want the graph view of the same idea, with decision traces and precedent, read what a context graph is and how it differs from a knowledge graph.
Frequently asked questions
What is a context layer in AI?
A context layer is the part of an AI stack between a company's data and tools and its AI models and agents. It connects to source systems, keeps facts current, enforces each user's permissions and assembles the relevant context for each request. The model stays the same; the input it receives gets better.
What is the difference between a context layer and a semantic layer?
A semantic layer defines what metrics mean and how they are calculated, so dashboards and queries agree. A context layer adds what an AI agent needs to act: who may see a fact, where it came from, whether it is still current, and what was decided in documents and conversations. If you run BI and want agents that act across tools, you need both, with the context layer reading metric definitions from the semantic layer.
Is RAG a context layer?
No. Retrieval-augmented generation is a technique for pulling relevant passages into a prompt for a single request. A context layer can use RAG, but it also handles state across tasks, entity resolution, permissions and deciding which facts are outdated.
What is an enterprise context layer?
An enterprise context layer is a context layer shared across a whole company rather than built for one agent or one team. It reads from many systems, including chat, email and meeting notes as well as databases, and serves the same governed context to every assistant and agent the company uses. Separate layers per team tend to drift apart into what Atlan calls context islands.
What is the purpose of a context layer for AI agents?
Agents act, so a wrong or outdated input becomes a wrong action rather than only a wrong answer. A context layer gives an agent the current facts, entities and permissions it needs at each step, so it does not have to rebuild that picture through a long chain of tool calls on every task.
Should we build or buy a context layer?
Build when the scope is narrow, when the data is mostly in a warehouse you already govern, or when the context layer is part of a product you sell. Buy when you need many SaaS sources, conversations and per-user permissions across several AI tools. Either way, the work sits in connectors, entity resolution, permissions, freshness and evaluation.
How do you test a context layer?
Use 50 to 100 real questions from your own work with agreed answers, including counting, list, negation and multi-hop questions. Add a permission-leak test with two users who have different access, and a stale-fact test in which a decision is reversed in a second tool. Then compare cost, speed and blind reviewer preference against your current setup.
Who owns the context layer?
DataWalk notes that ownership is often split across data engineering, governance, architecture and AI teams, with no single function responsible for the whole. A workable model is one accountable owner for the layer, such as the platform or data team, with each source system's owner responsible for its permissions and a security review for any change to what people can see.
Related reading
- What Is an Enterprise Knowledge Graph? Structure, Uses, and Limits
- Context Engineering: Deciding What an AI System Actually Sees
- What Is Contextual Understanding? A Guide to AI Interactions
- RAG vs MCP: What They Do, Why They Are Not Alternatives
- What Is AI Agent Memory? Types, Architectures, and How to Choose
- What Is MCP (Model Context Protocol)? A Complete 2026 Guide
- MCP Security: The Risks, and What Actually Mitigates Them
- What Is Organizational Memory?
- Introducing OM2: Your enterprise just started thinking
- Coworker vs Glean
- Benchmarks: OM2 vs native connectors
Ready to get started?
Put Coworker to work inside your actual stack
Connect Salesforce, Slack, Jira, whatever you use, and run your first agent in minutes.