On this page
Put Coworker to work on your stack.
Connect Salesforce, Slack, Jira and run your first agent in minutes.
Enterprise AI
What Is a Context Graph? How It Differs From a Knowledge Graph, and How to Build One
Coworker AI explains the context graph: a knowledge graph plus decision traces and time, how it differs from a knowledge graph, and how to build one.
A context graph is a knowledge graph of a company's entities and facts, extended with decision traces and time, that AI agents query for precedent and current state. A knowledge graph can tell an agent that Acme is a customer and who owns the account; a context graph can also tell it why Acme got a 20% discount when policy caps renewals at 10%, who approved the exception, and whether that precedent still applies.
The term went mainstream after Jaya Gupta and Ashu Garg of Foundation Capital published AI's trillion-dollar opportunity: Context graphs on December 22, 2025. Their argument is that systems of record store outcomes (the final price, the closed ticket) while the reasoning behind them lives in Slack threads, calls and people's heads, and that agents working inside a workflow can capture that reasoning as decision traces. Ahrefs estimates that US searches for "context graph" went from 19 in November 2025 to 3,244 in January 2026, and they were still about 776 a month in September 2026. Graph database vendors, enterprise search companies and agent memory projects now publish their own definitions.
A disclosure before going further: I work at Coworker, whose CEO describes its OM2 memory layer as a context graph. I use OM2 as the worked example near the end, next to Neo4j, Glean, Zep and Mem0, and I say where each of those is the better choice. Every vendor fact below comes from that vendor's own pages or papers, read on October 6, 2026.
What is a context graph?
A context graph connects three things that usually live in different systems: what a company knows (customers, projects, policies, people), what happened (messages, meetings, tickets, approvals), and the decisions that came out of it, with the evidence behind them. Each fact carries a time and a source, so an agent can ask what was true when a decision was made and also what is true today.
Vendors stress different parts of that. Neo4j describes three kinds of memory in one graph: long-term enterprise knowledge, short-term conversation history, and reasoning memory that holds decision traces. Graphwise describes a knowledge graph extended with time and decision lineage, where each statement carries metadata such as its validity period, source, confidence and access scope. Glean defines it as a model linking enterprise entities (people, documents, tickets, systems) to time-stamped traces of the actions and events between them. The shared core is entities, time, and a record of decisions.
What is a decision trace?
A decision trace is a structured record of one decision: the inputs gathered, the policy applied, any exception granted, who approved it, and what happened. Foundation Capital contrasts it with a rule. A rule tells an agent what should happen in general, such as which ARR definition to use in reporting. A decision trace records what happened in one specific case, including the policy version, the exception and the precedent relied on.
The essay's example is a renewal. An agent proposes a 20% discount against a policy that caps renewals at 10% unless a service-impact exception is approved. It pulls three SEV-1 incidents from PagerDuty, an open escalation in Zendesk, and an earlier thread where a VP approved a similar exception. Finance approves. The CRM ends up with a single value, the 20% discount, and none of the reasons for it.
Neo4j's guide to decision traces lists the same parts in engineering terms: the decision, the outcome, the reasoning steps, the tool calls with their results, and the context (entities and conversations) linked to each step. It also separates a decision trace from an LLM trace, which covers a single model run and sits in an observability tool, and from an application log, which records events and errors. Only the decision trace persists as memory the agent can query on the next case.
Where did the term "context graph" come from?
Foundation Capital popularized the term rather than coining it. Graphwise notes that the semantic web community spent more than a decade attaching context to knowledge graph statements, for example with contextualized knowledge repositories organized by time, space or topic. Ahrefs data shows the phrase drew between 5 and 37 US searches a month from January to November 2025, before the essay.
What the essay added was a market thesis. Gupta and Garg argue that the next large software platforms could be systems of record for decisions, and that startups running agents inside workflows are better placed to capture decision traces than incumbents built around current-state data. Three of the startups the essay uses as examples, Regie, Maximor and PlayerZero, are Foundation Capital portfolio companies, as the firm's one-month follow-up states. That does not make the argument wrong, and it is worth knowing when you weigh the claim that incumbents cannot build this.
The follow-up, posted January 30, 2026, also has the plainest definition I have found: a context graph is institutional memory for how an organization actually makes decisions, which is often different from what the process document says.
Does IBM mean the same thing by "context graph"?
No. IBM uses the term for a context engineering technique: structuring the information placed in a model's context window as a graph of nodes and edges, then selecting and ranking a subgraph to build the prompt. IBM's article names that pipeline GraphRAG, describes its graph as a representation of current state, and does not mention decision traces, which is close to the opposite emphasis from Foundation Capital's. The two meanings share one idea, using graph structure to choose what a model sees, and differ on whether the graph is a lasting record of the company. For the broader practice IBM is describing, see this guide to context engineering.
Context graph vs knowledge graph vs vector database vs context layer
These four terms get used interchangeably in vendor copy. They describe different layers, and most working systems combine them.
| Knowledge graph | Context graph | Vector database | Context layer | |
|---|---|---|---|---|
| What it is | Entities and typed relationships with shared business meaning | A knowledge graph plus decision traces, timestamps and sources | An index of embeddings for similarity search | Everything that decides what a model sees for a request: retrieval, memory, ranking, permissions, delivery |
| What it returns | Entities, relationships and paths | Facts with validity windows, sources and linked decisions | The k most similar chunks of text | An assembled context for one task |
| How it handles time | Usually stores the current state | Records when each fact became true and when it was superseded | Only through timestamp metadata you add | Inherits whatever its stores support |
| A question it answers well | "Who owns the Acme account?" | "Why did Acme get 20% when the cap is 10%, and does that precedent still apply?" | "Find tickets that sound like this complaint" | "Give this agent what it needs for the Acme renewal, and nothing this user cannot see" |
| Where it struggles | Why something happened, and change over time | Capturing intent; it needs instrumentation where decisions are made | Counting, listing every match, multi-hop questions | Parts built separately drift apart |
| Examples | Neo4j, Graphwise GraphDB | Graphiti, Neo4j Agent Memory, Glean, OM2 | Pinecone, vector indexes in graph databases | Zep, Glean, OM2 |
Sources: Neo4j's graph, knowledge graph and context graph comparison, Graphwise, Pinecone's definition of a vector database and Zep, each read on October 6, 2026. The example questions are mine.
Context graph vs knowledge graph: what is the difference?
A context graph is a knowledge graph with time, provenance and decisions added. Neo4j frames it as a ladder: a graph shows what is connected, a knowledge graph adds what those connections mean, and a context graph applies that knowledge to a task, answering which context matters right now and why. In Neo4j's model the knowledge graph is the agent's long-term memory, and the context graph connects it to conversation history and decision traces.
Graphwise has the clearest single example. Ask a knowledge graph who Alice works for and it returns Company X. Ask a context graph who employed Alice on January 15, 2026, and it returns the answer for that date, plus the fact that it came from HR records, carries 95% confidence, and is visible only to finance and compliance roles. Time, source, confidence and access are all attached to the one fact.
The hard problems of the knowledge graph underneath (entity resolution, stale data, permissions, extraction quality) are covered in the separate guide What is an enterprise knowledge graph?. All of that still applies. A context graph adds more things to get wrong.
Where does a vector database fit?
A vector database indexes embeddings so you can find text that means something close to a query. It is the right tool for "find passages like this one," and it usually sits inside a context graph system rather than competing with it. Graphiti, for example, combines semantic embeddings, BM25 keyword search and graph traversal in its hybrid retrieval.
What similarity search cannot do on its own is answer questions about structure. Top-k retrieval returns the k closest chunks, so "how many open escalations mention SSO?" or "list every customer who asked for SSO" comes back incomplete by design. IBM makes a related point: context chosen by top-k retrieval alone misses indirect relationships. The GraphRAG paper behind Microsoft's GraphRAG project found that plain RAG fails on questions about a whole corpus, such as the main themes in a dataset, and that its graph-based approach gave substantially more comprehensive and diverse answers on datasets of around 1 million tokens. The RAG vs MCP explainer covers where retrieval stops and tool calls take over.
Where does the context layer fit?
The context layer is the broader system that decides what a model sees for a given request: retrieval, memory, ranking, permission checks and delivery to the AI tool. A context graph is one store inside it, usually the one holding relationships and history. Glean's engineers draw the same boundary from the other direction. They argue that high-value workflows need both a context layer with a process-aware model of the company and an execution layer that plans and generates traces, and that separating the two lets them drift apart. The layer exists because of two limits: a context window only holds so much, and context rot means accuracy drops as that window fills. A company does not fit in a prompt.
Why do AI agents need a context graph?
Agents fail in the gap between what systems store and what people know. In a hands-on post for Neo4j, William Lyon names the two views. The state clock is what is true right now, the way a database knows a customer's credit limit is $50,000. The event clock is what happened, when, and why. Traditional databases, as Lyon points out, are built for the state clock. An agent asked to make a judgment call needs the event clock as well.
Precedent: how were similar cases handled?
Exception-heavy work runs on precedent. Foundation Capital lists the knowledge that never reaches a system of record: exception logic that lives in people's heads (its example is an extra 10% discount for healthcare customers because of their procurement cycles), precedent from past deals, synthesis across systems, and approvals given on a call or in a direct message. The essay points to deal desks, underwriting, compliance reviews and escalation management as the places where this matters most. Without precedent, an agent has two options: send every case to a human, or guess.
Current state: is this still true?
The opposite failure is acting on a fact that used to be true. Graphwise describes it as an agent retrieving a policy or fact that is no longer valid and acting on it with high confidence. A context graph addresses it by storing when each fact became true and when it stopped, so both "what is the current discount policy?" and "what was the policy when this deal closed?" have answers. Garg's follow-up names the hard part. Decisions have a half-life, and a pricing exception approved three years ago under a different CFO may not apply today.
Structured questions: list it, count it, follow the chain
Many real questions are not lookups. "Which enterprise customers asked for SSO this quarter?" needs enumeration. "How many P1 incidents touched billing in September?" needs counting. "Which open deals did the account executive who left in August own, and who owns them now?" needs several hops across people, deals and ownership changes. A graph answers these with a query, while similarity search answers them with a sample. In the Neo4j post, tracing which decisions led to an account freeze is a two-line Cypher query that follows CAUSED relationships up to five steps back, where SQL would need recursive common table expressions and self-joins.
How do you build a context graph?
Building one is mostly plumbing and policy rather than modeling. The seven steps below cover what an implementation has to handle, whether the code is yours or a vendor's.
1. Connect the tools where decisions happen
As Foundation Capital's essay argues, decisions mostly happen outside the system of record. They happen in Slack or Microsoft Teams threads, in meetings recorded through Zoom, Gong or Granola, in Gmail or Outlook, in Jira comments, in CRM notes and in shared docs, and the system of record receives the outcome. So the first job is connectors that read three things from each tool: content, activity (edits, comments, approvals, status changes) and permissions.
Glean, which says it spent six years building its context platform, describes moving from indexing documents to capturing every change event in an app and normalizing those events into traces, with the aim of building its context graph. Its CEO gives a concrete example: a Salesforce connector may expose a deal-stage change, but understanding that change also takes the Google Doc edit, the Slack message and the calendar event around it.
2. Resolve entities across tools
The same customer is "ACME Inc" in the CRM, "ACME" in support tickets and "acme-prod" in an incident channel. Until those resolve to one node, the graph holds three partial customers. Glean describes running a machine learning pipeline over content and activity signals to infer projects, customers, products, teams and people, and uses the ACME case as its example. People are a hard case of their own: one person can have a Slack handle, two email addresses, a GitHub username and a CRM owner ID. Errors here compound, because every fact attached to the wrong node is wrong in a way no prompt can fix.
If you are building on open source, LLM-based extraction gives you candidate entities quickly. Neo4j's LLM Knowledge Graph Builder turns PDFs, documents, web pages and YouTube transcripts into a graph of text chunks plus an entity graph, using models from OpenAI, Google (Gemini) and Anthropic (Claude), among others, with an extraction schema you can configure. Matching those candidates to the people and accounts already in your systems is still work you have to design.
3. Capture decision traces
This step is what makes it a context graph rather than a knowledge graph, and traces come from two places. The first is agents. When an agent runs a workflow, it can record the inputs it gathered, the policy it checked, the exception path it took and the result. Foundation Capital argues this is where the advantage sits, because the orchestration layer sees the full context at the moment of decision rather than after an ETL job. Neo4j's decision traces guide makes the engineering version of the point: with its Agent Memory library, trace capture is explicit, and your code decides what gets recorded and when.
The second source is people. Many decisions are still made by humans in conversation, so a context graph also has to read the Slack thread where an exception was argued and the meeting where it was approved. Glean's CEO is candid about the limit here: the reasoning behind a decision often stays in someone's head, while the process (steps, approvals, field changes) leaves a trail that can be modeled.
If you design the schema yourself, the W3C PROV-O ontology is a sensible starting point. Its three core classes, Entity, Activity and Agent, map onto a decision record (the thing decided, the act of deciding, and the person or agent responsible), and properties such as used, wasGeneratedBy and wasDerivedFrom express provenance chains. The Neo4j hands-on model takes a similar shape, with Decision, Exception and Escalation nodes and relationships such as APPLIED_POLICY, GRANTED_EXCEPTION and PRECEDENT_FOR.
4. Model time and retire outdated facts
Every fact needs at least two timestamps: when it became true and when it stopped being true. Graphiti, Zep's open-source framework, stores each fact as an edge with a validity window. When new information contradicts an old fact, the old one is invalidated rather than deleted, so you can ask what is true now or what was true on any past date. Graphwise uses W3C standards such as RDF-Star and OWL-Time for the same purpose.
Storage is the easy half. The hard half is detecting that a fact has been superseded. A pricing policy on a Confluence page can be overridden by a message from the CFO in Slack, and nothing in either system links the two. A decision made in a meeting can be reversed in a later thread. Garg's follow-up lists handling time as an open problem, with no agreed rule for when a past decision stops applying. Any vendor that claims to handle this should be able to show you, on your own data, a fact that changed and the graph noticing.
5. Carry permissions through every fact
This is the step that decides whether your security team will let you ship, and the one explainers cover least. Of the five context graph explainers I read for this guide (Foundation Capital, IBM, Neo4j, Graphwise and Glean), two never mention permissions, and only Graphwise attaches access rules to individual facts. None of the five walks through what happens to a fact derived from sources with different access rules.
A context graph is built from sources with different rules: a private Slack channel, an HR folder, a customer record in Salesforce. Once those are broken into facts and connected, the graph can answer questions that cross the boundaries between them. The rule I would hold any implementation to is simple to state: a fact is visible only to people who can see its source, and the check runs on every traversal, not once at ingestion. A summary built from a public channel and a private one should be visible only to people who can see both, or be split into separate facts.
The checks also have to be fast, because a multi-hop query can touch thousands of edges. Google's Zanzibar paper shows what authorization looks like at that scale. The system stores and evaluates access control lists for hundreds of Google services, including Calendar, Drive and YouTube, handles millions of authorization requests per second, and held 95th-percentile latency under 10 milliseconds over three years in production. If you build on an agent memory library, check what its identifiers actually enforce. Neo4j's Agent Memory documentation says its conversation IDs identify conversations rather than enforce access control.
6. Support structured queries, not only search
Agents ask two kinds of questions: find me something, and tell me something exact. The second kind needs queries or typed tools over the graph that can enumerate every match, count, confirm that something did not happen, rank by structure, and follow relationships across several hops. On a property graph the usual language is Cypher, Neo4j's declarative query language. Rather than asking the model to write raw queries, Neo4j's own demo gives a Claude agent 11 purpose-built MCP tools for its context graph, including tools to find precedents, trace a causal chain and record a decision.
7. Deliver context to agents over MCP
The last step is putting the graph inside the AI tools people already use. The Model Context Protocol is an open standard for connecting AI applications to data and tools, supported by clients including Claude, ChatGPT, Visual Studio Code and Cursor. Neo4j Agent Memory, Mem0 and OM2 can all be reached through MCP servers. For background, see what MCP is.
The design question that matters most for a context graph is whose permissions apply. The MCP authorization specification makes authorization optional. When an HTTP-based server implements it, the server acts as an OAuth 2.1 resource server and must reject access tokens that were not issued for it. For company data, insist on per-user authorization so the graph answers as that person, rather than through one service account that can see everything. The guides to MCP security and enterprise MCP cover the common failure modes.
Coworker
Answers from all your company knowledge
Coworker connects your tools and acts on what it finds, not just search.
Ask Coworker across your toolsShould you build or buy a context graph?
Both are reasonable. The answer depends mostly on where your decisions are made and who has to maintain the result.
When to build on Neo4j, Graphiti or similar
Building makes sense in three situations:
- The context graph is part of your own product or a core workflow you run, such as a credit decision system or a claims process, and you need control over the schema, the decision record and where the data lives.
- Your decisions already flow through software you control, so you can capture traces at the moment of decision. That is the case Foundation Capital makes for agent startups.
- You have engineers to run a graph database, extraction pipelines and evaluation for years, not weeks.
The parts exist. Neo4j and Cypher handle storage and queries, Graphiti adds validity windows and provenance (it runs on Neo4j, FalkorDB or Amazon Neptune), PROV-O gives you a decision schema, and MCP handles delivery.
When to buy a platform
Buying makes more sense when the decisions you care about are made by people across Slack, meetings, email, Jira and CRM, and the hard part is connectors, entity resolution and permissions across many tools rather than modeling one workflow. It also fits when you want the same company context inside several AI tools at once (Claude for most staff, Cursor for engineering, ChatGPT for a few teams) without building and securing your own MCP server, and when security will not approve anything that cannot show inherited source permissions and an audit trail. Even then, keep building decision records for the exception-heavy workflows you automate yourself, because a platform that reads your tools does not sit inside your own agent's approval step.
| Component | Build on Neo4j or Graphiti | Buy a platform |
|---|---|---|
| Connectors and change events | You write and maintain one per tool | Included; check content, activity and permission depth per tool |
| Entity resolution | You design matching rules and review queues | Included; test it with your messiest customer names |
| Decision traces | You instrument your own agents and workflows | Varies: process patterns, facts from conversations, or both |
| Time and supersession | Graphiti has validity windows; on plain Neo4j you model it yourself | Ask to see a reversed decision handled on your data |
| Permissions | You sync access lists and check them on every query | Should be inherited from each source and enforced on traversal |
| Delivery to agents | You run an MCP server and OAuth | Usually an MCP server plus the vendor's own apps |
How do you test a context graph before trusting it?
Whichever path you take, test on your own data. Vendor benchmarks, including the one I cite below, use someone else's questions. Five tests catch most problems:
- Golden questions: 30 to 50 real questions your team asked last month, with answers checked by the people who know. Include enumeration, counting and "why" questions, not only lookups.
- Permission test: two people with different access ask the same question. The answers should differ exactly as their access to the sources does.
- Stale-fact test: pick a decision that was later reversed and check that the graph returns the current one, and can show the older one with dates.
- Provenance test: every answer should point to the message, document or record it came from.
- Cost and speed: tokens and time per answer, measured against your current setup.
Context graph platforms in 2026: who offers what
Here is how the main options compare, using each vendor's own pages as of October 6, 2026. They are not all the same kind of product, which is the first thing to understand before comparing them.
| Platform | What it is | How it approaches the context graph | Best fit |
|---|---|---|---|
| Neo4j | Graph database, plus the Agent Memory library from Neo4j Labs | Long-term knowledge, conversation history and reasoning traces in one graph | Teams building their own context graph |
| Glean | Enterprise search and agent platform | Change events across apps, aggregated into anonymized process patterns | Large companies standardizing on Glean for search and agents |
| Zep and Graphiti | Managed context platform, and its open-source temporal graph framework | Facts with validity windows, raw episodes as provenance, a graph per user, account or domain | Developers building agents where facts change over time |
| Mem0 | Memory layer for AI agents and apps | Entities link the memories that mention them, without typed relationships | Per-user memory inside an AI product |
| Coworker OM2 | Organizational memory across a company's tools | Atomic facts, entities and relationships from 50+ connectors, served over MCP and native apps | Company-wide context inside Claude, ChatGPT, Cursor and other clients |
Neo4j: the graph database you build on
Neo4j is a graph database you can build a context graph on yourself, and its Agent Memory library stores the three memory types in one graph. Neo4j lists five ways to start: a hosted Agent Memory Service with an API key, your own Neo4j deployment (Aura, Desktop or Docker, including air-gapped use), an MCP server, a create-context-graph scaffold with a FastAPI backend and a Next.js frontend, and Aura Agent, which is in early access. Two things to know before you commit. Agent Memory is a Neo4j Labs project, actively maintained but not officially supported, with no SLAs or backward-compatibility guarantees. And as noted above, its conversation IDs do not enforce access control. It is the best fit for teams that want to own the schema and have engineers to run it.
Glean: context graphs from process traces
Glean's CEO, Arvind Jain, welcomed the term in January 2026, writing that the company was glad the idea finally had a name and that Glean has its context platform in place, including the graph, after six years of technical investment. Its engineering post treats actions (created, viewed, approved, escalated, resolved) as nodes with timestamps, and treats edges as causality and correlation with a probability attached. That gives agents a distribution of likely paths instead of a hard-coded flow. Underneath sit document-level connectors, an inferred knowledge graph with entity resolution, and a personal graph visible only to its owner. Aggregate patterns are built from anonymized steps and count only if they appear across a minimum number of distinct users and traces.
Glean's strengths are real: connector breadth (it says it offers more than 250 connectors, and its connectors page lists 275+), permission enforcement, and a privacy design with explicit anonymity thresholds. Its own stated limits are worth reading too. The CEO puts its task understanding at about 80% accuracy, the engineers say their storage model is not optimized for reasoning across thousands of incidents at once, and Glean says long-running autonomous agents are close but not yet a reality. The design difference that matters is scope. Glean models how work typically gets done across many people, while Foundation Capital's definition centers on specific decision records. Learning how P1 incidents usually get resolved is Glean's design center; finding out who approved one particular exception, and under which policy, needs record-level traces. For a direct comparison, see Coworker vs Glean and the Glean Agent Builder guide.
Zep and Graphiti: temporal context graphs for agents
Graphiti is an open-source framework (Apache-2.0, about 31,500 GitHub stars) for building temporal context graphs. Entities are nodes, facts are relationships with validity windows, and every derived fact traces back to the raw episode that produced it. Retrieval combines semantic, keyword and graph search, and it runs on Neo4j, FalkorDB or Amazon Neptune. Zep is the managed service built on it, now positioned as a context layer for enterprise data: it builds a context graph per user, customer, account or domain, and applies policies to decide what each agent may retrieve. Zep's paper on arXiv reports 94.8% on the Deep Memory Retrieval benchmark against 93.4% for MemGPT, and accuracy gains of up to 18.5% with 90% lower response latency than baseline implementations on LongMemEval. It fits developers who need agent memory where facts change and history matters.
Mem0: memory for agents and apps
Mem0 describes itself as a memory layer for AI agents and apps. Its platform builds a graph automatically: entities mentioned in memories become nodes, and memories that share an entity are linked, with no separate graph database to run. Mem0's graph memory documentation is explicit that it does not assign typed relationships between entities (it will not record that one person manages another), because connections come from co-occurrence. Mem0's paper reports a 26% relative improvement over OpenAI on an LLM-as-a-judge metric on the LOCOMO benchmark, a graph variant scoring about 2% higher overall than the base version, and 91% lower p95 latency with more than 90% token savings compared with sending the full conversation. It fits per-user personalization inside an AI product. The Mem0 vs Zep vs Letta comparison and the guide to AI agent memory go deeper.
OM2: company-wide memory from the tools you already use
OM2 is organizational memory that reads a company's connected tools and serves the result to AI clients over MCP and to the native apps built on it. It is built for the buy case above, where decisions are spread across many tools and several AI clients need the same context with source permissions intact. The next section walks through it as a worked example.
Worked example: how OM2 works as a context graph
The OM2 launch post by Alex Calder, the CEO, says it directly: "You can think of OM2 as a context graph." The post goes on to call it a neural graph, a name it uses because of the graph's density, to set it apart from document-level knowledge graphs. Here is how the public description maps onto the build steps above.
| Build step | What the public pages say OM2 does | Source |
|---|---|---|
| Connect sources | Reads from 50+ connectors (including Slack, Gmail, Zoom, Gong, Jira, Salesforce, HubSpot, Google Drive, Notion and Confluence) continuously, with no uploads or re-syncing | Connectors, Platform |
| Resolve entities | Discovers people, companies, projects and deals across connected apps and maps relationships from day one | OM2 page |
| Facts and decisions | Breaks documents, messages and tickets into atomic facts enriched with entities and relationships, stored as connected facts about people, projects, customers and decisions | OM2 page, Platform |
| Time | Learns backwards over historical data at setup, updates throughout the day, and invalidates outdated information | Launch post |
| Permissions | Access policies inherited from the underlying sources travel with every fact and connection; source permissions are enforced on traversal, with provenance chains back to the sources | Launch post |
| Queries | Structured queries for enumeration, negation, counting, structural ranking and multi-hop traversal, with context ranked by proximity, recency, entity centrality and relationship strength | Launch post, OM2 page |
| Delivery | Connects via MCP or API, so it works in Claude, ChatGPT, Gemini, Perplexity and custom agents; the MCP page also lists Claude Code, Cursor and Windsurf | Launch post, Coworker MCP |
Two practical consequences follow from that design. Two people asking the same question can get different answers, matching what each can see in the source tools. And because OM2 builds itself from the tools a company already uses, there is no schema to design and no migration.
It is also worth being clear about what OM2 is not. It is not a graph database you model yourself. If you are building a decision-record system inside one workflow you own, Neo4j or Graphiti is the better fit, and if you need per-user memory inside your own consumer app, Mem0 or Zep is. Formal decision records with policy versions and named approvers, the kind Foundation Capital describes, still come from instrumenting the workflow where a decision is made; OM2's public description centers on connected facts drawn from the tools where decisions get discussed. The public pages say OM2 invalidates outdated information without going into how a superseded fact is detected, so the stale-fact test above is worth running on your own data.
On cost and speed, the published benchmark covers 100 tasks across seven categories, run on a Claude harness against nine connected sources: Google Drive, Gmail, Google Calendar, HubSpot, GitHub, Jira, Slack, Notion and Stripe. It compared Claude using OM2 over MCP with Claude using its own native connectors. In the best case, with OM2's Learning feature (which improves retrieval as a team uses it) on the retrieval-heavy task categories, OM2 delivered 9x lower token cost (89.1%) and finished tasks 64% faster. Across all seven categories the same setup was 75.5% cheaper and 45.6% faster. In blind pairwise comparisons with the baseline, judged against a human-authored reference answer with the order randomized, reviewers preferred the OM2 answers 84.5% of the time.
Statistically significant benchmarks comparing Coworker MCP vs. Claude Native Tooling.
On security, the launch post lists SOC 2 Type II, GDPR and CASA Tier 2, and says customer data is never used to train models. Background on the category is in What is organizational memory in AI?.
What are the limits of a context graph?
A context graph is useful, and it is also new enough that the hard problems are still open. Four are worth knowing before you invest.
You cannot fully capture the why
According to Foundation Capital's own follow-up, the most common objection to its essay was that intent is internal and unobservable, so what you can reliably capture is the sequence of actions. Glean's CEO made the same argument. Garg agreed with a qualification: record the policy applied, the evidence consulted, the exception granted and the approver, then infer the reasoning from patterns over time. That is a useful frame for buyers. A context graph increases how much of a decision gets recorded, and it does not read minds.
Precedent ages out
A graph that keeps every decision forever will eventually hand an agent a precedent set by people who have since left, under a policy that has since changed. Someone has to decide when a precedent stops counting: after a policy version changes, after a reorganization, or after a fixed period. Foundation Capital lists this as an open question. In my view it is a governance decision more than a technical one, and it belongs with whoever owns the policy.
Decision traces are sensitive data
A record of who approved what, and why, is the kind of data legal and HR teams will ask about. Garg's follow-up asks who gets to query decision traces that include sensitive reasoning about customers, employees or risks, and how long they should be kept. Glean's answer for its aggregate graph is to leave out raw text and user identifiers and keep only patterns seen across many users. Whatever you choose, settle retention and access for traces before you start collecting them.
It may end up a feature rather than a category
Garg reports that some skeptics see context graphs as "data catalog 3.0," infrastructure that ends up inside warehouses, catalogs and observability tools rather than becoming a product category of its own. His follow-up also expects every organization to run several context graphs, each shaped by its domain. For a buyer, the category debate matters less than whether a graph answers your questions on your data, with your permissions. The best enterprise AI knowledge management platforms roundup shows how differently vendors package the same underlying ideas, and the guide to contextual understanding in AI covers why models need this context in the first place.
If you want to see a context graph built from tools like yours, book a demo of OM2. If you would rather start with numbers, ask for a Context Impact Report: Coworker connects your sources, builds OM2 on your data, runs the benchmark questions against your current setup and through OM2, scores every answer blind, and covers the token costs.
For where a context graph sits in the wider AI stack, and how it compares with a semantic layer, see what a context layer is.
Frequently asked questions
What is a context graph in AI?
A context graph is a knowledge graph of a company's entities and facts, extended with decision traces and time, that AI agents query for precedent and current state. It records what happened and why, with dates and sources, so an agent can check how similar cases were handled and whether a fact is still true. Foundation Capital popularized the term in a December 2025 essay.
What is the difference between a context graph and a knowledge graph?
A knowledge graph stores entities and the relationships between them, usually as a current snapshot. A context graph adds time (when each fact was true), provenance (where it came from) and decision traces (what was decided, by whom, under which policy). A knowledge graph can say who owns a customer account today; a context graph can also say who owned it last March and why that customer's discount was approved.
Who coined the term "context graph"?
Jaya Gupta and Ashu Garg of Foundation Capital popularized it in their December 22, 2025 essay, "AI's trillion-dollar opportunity: Context graphs." The phrase existed earlier in semantic web research on adding context to knowledge graph statements, but it drew fewer than 40 US searches a month before the essay. IBM uses the same term differently, for structuring a model's context window as a graph.
What is a decision trace?
A decision trace is a structured record of a single decision: the inputs gathered, the policy applied, any exception granted, who approved it, and the outcome. Neo4j's version also includes the reasoning steps and tool calls an agent made along the way. Traces accumulate into a context graph, which lets agents search past decisions as precedent.
How do you build a context graph?
Connect the tools where decisions happen (chat, meetings, email, tickets, CRM and docs), resolve the same people and accounts across those tools, and record decisions with timestamps and sources. Then retire facts that are no longer true, enforce each source's permissions on every query, support structured queries such as counting and multi-hop traversal, and deliver the result to agents, usually over MCP. Teams either build on a graph database such as Neo4j or a framework such as Graphiti, or buy a platform that handles connectors and permissions.
Can a context graph capture why a decision was made?
Only partly. Intent often stays in someone's head, so most systems capture the observable parts of a decision (the evidence consulted, the policy applied, the exception granted and the approver) and infer the reasoning from patterns across many cases. Foundation Capital and Glean both describe the limit this way, and Foundation Capital calls it the most common pushback its essay received.
What does Glean mean by a context graph?
Glean defines a context graph as a model connecting enterprise entities such as people, documents, tickets and systems with time-stamped traces of the actions between them. It builds the graph from change events across its connectors, aggregates anonymized steps from many users into likely process paths, and uses those paths to guide agents. Its focus is how work typically gets done, rather than a record of each individual decision.
Is a context graph the same as AI agent memory?
They overlap but are not the same. Agent memory usually means what one agent or app remembers about its users and sessions, which is what libraries like Mem0 are built for. A context graph is broader: shared, company-wide knowledge plus decision history that many agents and people query, with permissions inherited from the source systems.
Related reading
- What Is an Enterprise Knowledge Graph?
- Introducing OM2: Your enterprise just started thinking
- Organizational Memory (OM2)
- What Is Organizational Memory in AI?
- What Is AI Agent Memory?
- Mem0 vs Zep vs Letta: Compared
- What Is MCP (Model Context Protocol)?
- RAG vs MCP: The Difference Explained
- Context Engineering: A Practical Guide
- Coworker vs Glean
- 8 Best Enterprise AI Knowledge Management Platforms
- AI Benchmarks: OM2 vs Native Connectors
Ready to get started?
Put Coworker to work inside your actual stack
Connect Salesforce, Slack, Jira, whatever you use, and run your first agent in minutes.