On this page

Put Coworker to work on your stack.

Connect Salesforce, Slack, Jira and run your first agent in minutes.

Book a demo
Blog

Enterprise AI

Generative Engine Optimization: How to Get Cited by AI Search

Coworker AI explains generative engine optimization: how AI answers pick sources, what actually moves citations, and how to measure it honestly.

Dhruv Kapadia18 min read

Generative engine optimization, usually shortened to GEO, is the practice of making a page likely to be selected, quoted, and attributed inside an AI-generated answer. The term comes from a 2023 research paper that framed the problem formally and tested which content changes increased a source's visibility in generated responses.

You will also see answer engine optimization (AEO) and AI search optimization used for roughly the same idea. The distinctions people draw between them are mostly marketing. The useful split is between optimizing to be clicked and optimizing to be cited, and that split is real.

Why this became a separate objective

For twenty years the goal was a click. A search returned links, and ranking higher produced more visits.

AI answers break that chain. The user asks a question, receives a synthesized answer drawing on several sources, and often has no reason to visit any of them. The page still did its job. The traffic report will not show it.

We measured this on our own property and published the study: across 906,663 non-branded impressions over 90 days, click-through rate fell as query length rose while average position improved. Queries of twelve words or more averaged position 5.5 and returned a single click from 61,782 impressions. Nearly 69% of non-branded impressions produced no click at all.

That is the case for GEO in one number. On the queries where our content ranked best, being clicked had almost stopped happening. Being the cited source was the only available return.

How do AI answers choose sources?

Nobody outside the labs has the ranking function, and anyone claiming otherwise is guessing. What can be established from documentation and observed behaviour:

Retrieval usually starts from search. Most AI answer surfaces retrieve candidate documents using an underlying search index before generating. That means classic SEO is the entry ticket. A page that cannot be found is not a candidate to be cited, and this is the single most common reason a GEO programme produces nothing: the content is fine and the pages were never in the candidate set to begin with.

Passages are selected, not pages. The unit is the passage that answers the question, not the document. A long page can be cited for one paragraph while the rest is ignored, which changes how you structure content.

Extractability matters more than persuasion. The model needs a clear, self-contained statement it can lift. Prose that builds to a conclusion over four paragraphs is harder to quote than prose that states the conclusion and then supports it.

Length is not the variable people think it is. Long pages are not favoured for being long. They tend to do well because they contain more distinct passages, each of which can answer a different question, which is a coverage effect rather than a word-count effect. Padding a page to hit a target does nothing; adding a genuinely new section that answers a question the page did not previously address does.

Attribution follows the original. A statistic is attributed to whoever produced it. Restating someone else's number makes you an intermediary the model has no reason to name.

Consensus is weighed. Claims corroborated across independent sources appear more often than contested ones. This makes being cited elsewhere part of being cited here, and it is the reason a GEO programme that only touches your own site tends to plateau. If three independent write-ups say the same thing and yours says something different, yours is the one that gets left out, whether or not you are right.

What does the independent evidence say?

The direction is well corroborated even though the mechanics are opaque.

Pew Research found users markedly less likely to click a link when an AI summary is present. Ahrefs measured a substantial decline in click-through for queries showing an AI Overview. Semrush's analysis found those surfaces concentrated on longer, question-shaped queries, which matches what we found by query length. And SparkToro's zero-click research documented the underlying trend well before AI answers arrived, which is a useful reminder that this did not start in 2024.

The historical click curves are worth knowing as a baseline, because the gap against them is the size of the change. Long-running studies from Advanced Web Ranking and Backlinko put position one somewhere around a quarter of clicks. Our own non-branded position 1-2 bucket returned 0.849%.

What the original research tested

The GEO paper is worth reading rather than taking second-hand, because most summaries of it overstate its findings. It built a benchmark of queries, generated answers, and tested which modifications to a source document increased its visibility in the generated output.

The changes that helped most were the unglamorous ones: adding citations to authoritative sources, including relevant statistics, and quoting credible references. Keyword stuffing did not help. That result has held up in practice and it is the empirical basis for the "publish specific, sourced, quotable claims" advice that follows.

Treat it as directional. It predates the current generation of AI search surfaces, it used a constructed benchmark rather than live production systems, and the providers have changed substantially since.

The technical layer

None of the content work matters if the machines cannot fetch and parse the page.

Crawl access

AI providers use their own crawlers, distinct from Googlebot, and each can be allowed or blocked independently in robots.txt. Blocking them is a legitimate business decision that some publishers have made deliberately. What is not legitimate is doing it by accident and then wondering why you are never cited. Check what your robots.txt actually says before starting a GEO programme.

There is a real tension here and it deserves stating plainly. Allowing AI crawlers means your content trains and grounds systems that may answer your buyers' questions without sending them to you. The trade is presence in the answer against traffic to the page. For most B2B companies presence wins, because the alternative is a competitor being named instead. For a publisher whose revenue is page views, the calculation is genuinely different.

Rendering

Content that only exists after client-side JavaScript execution is a risk. Search engines largely handle it; AI crawlers are less consistent. Server-rendered HTML is the safe default, and it is one of the few areas where the right answer has not changed in a decade.

Structured data

Schema.org markup does not directly cause citations, and claims that it does are unsupported. What it does is remove ambiguity about what a passage is, who wrote it, and when it was published. Author and date fields matter more than they used to, because recency and provenance are both weighed.

Site structure

Definitional content compounds when it sits together and interlinks. A cluster of related terms that reference each other gives a retrieval system multiple entry points into the same topic, and gives you internal links to distribute. This is why glossaries and reference sections outperform the same content scattered across an undifferentiated blog. Our own context window and agent memory entries exist as a cluster for exactly this reason.

Common mistakes

Patterns worth avoiding, several of which we have made.

Treating GEO as a separate content type. It is a different objective for the content you already need. Companies that spin up a parallel "AI content" workstream usually produce thin material that serves neither goal.

Chasing the head term. "Generative engine optimization" has meaningful volume and is heavily contested. The definitional long tail around your actual category is cheaper, more relevant, and more likely to be reached for when a buyer asks a real question.

Optimizing pages that were never going to be cited. A pricing page will not be quoted as an authority on a concept. Leave it alone and let it convert.

Publishing restated statistics. If you cite someone else's figure, they get the attribution. This is the difference between contributing to the conversation and being the reason it happens.

Declaring victory on noisy data. Citation counts move week to week for reasons that have nothing to do with your content. A single good week is not a result.

Ignoring the commercial half. The most expensive version of this mistake is redirecting an entire content operation toward citations while the pages that actually produce pipeline go stale.

What actually moves citations

Ranked roughly by how much they matter, based on what we have observed on our own content.

Publish something only you can publish

The single biggest lever, and the most underused. Original data, original benchmarks, original methodology. A model summarizing a topic names the source of a figure, so producing figures makes you nameable.

This does not require a research department. It requires looking at data you already have. Our study came from Search Console access we already had and a cut nobody had published.

Answer in the first two sentences

State the answer, then explain it. The inverted pyramid was always good practice for readers, and it is close to mandatory for machine extraction. If a model must read six paragraphs to work out your position, a competitor who stated theirs immediately is easier to quote.

Write headings as the questions people actually ask

Question-shaped H2s and H3s let a parser map a query to a section. "How much does it cost" beats "Pricing considerations", because the former matches the shape of the query.

Use real structure, not the appearance of structure

Genuine HTML tables, genuine lists, and structured data on question-shaped sections. Pipe characters arranged to look like a table in plain text are invisible to a parser. This is a common and entirely avoidable failure.

Be specific enough to be worth quoting

"AI reduces operational costs" attributes to nobody. "Queries of twelve or more words returned one click from 61,782 impressions" has an owner. Specificity is what survives summarization.

State your limitations

Counterintuitive but consistent in our experience: pages that say what their data cannot show read as more reliable. It also protects you when a reader checks, which some will.

Keep facts current and dated

Models weight recency for anything that changes. Pricing, product capabilities and market figures need a visible date and real maintenance. A stale number cited widely is worse than no number.

Earn mentions elsewhere

Being referenced across independent sites raises the odds of being selected. This is the part GEO shares most with traditional digital PR, and it is why third-party listicles and comparison roundups matter more than their referral traffic suggests.

Coworker

Put Coworker to work on your actual stack

Connect Salesforce, Slack, Jira and run your first agent in minutes.

Book a demo

What about llms.txt?

llms.txt is a proposed convention: a markdown file at your domain root offering a clean, structured summary of your site for language models, in the spirit of robots.txt but for comprehension rather than permission.

Worth doing, with realistic expectations. It is a proposal, not a standard, and adoption by major AI providers is not confirmed. It costs very little to publish and maintain. It is not a ranking lever and anybody selling it as one is overstating the case.

We publish one. We do not attribute any of our citation growth to it, because we have no way to measure that it did anything.

GEO versus SEO: what actually differs

Traditional SEOGenerative engine optimization
ObjectiveRank, earn the clickBe selected and quoted as a source
UnitThe pageThe passage
Success signalClicks, sessions, conversionsCitations, mentions, impressions
Content shapeComprehensive, keyword-alignedExtractable, specific, well-attributed
FreshnessMatters for some queriesMatters for most factual claims
Off-siteLinksLinks plus mentions and corroboration
MeasurementMature, directImmature, sampled, indirect

The overlap is larger than the difference. Crawlability, information architecture, internal linking, page speed and topical depth serve both. GEO is not a reason to build a separate content operation. It is a reason to change what you optimize for and how you judge the result.

How do you measure GEO?

This is where the discipline is weakest and where most claims fall apart.

Citation tracking tools sample a set of prompts and record which domains get cited. They are useful for direction and unreliable as a census. Different tools disagree, prompt sets differ, and results move as models update. We track ours weekly and read the trend, not the absolute number.

Impressions in Search Console become a proxy for the surfaces Google controls, since a page cited in an AI Overview accrues impressions. Rising impressions with flat clicks on question-shaped queries is a signal worth watching rather than a failure to fix.

Branded search should, in theory, rise if a brand is being surfaced repeatedly during research. In our own data branded search fell over a period when citations rose sharply, so we do not treat it as a reliable proxy.

Direct prompting is crude and honest: ask the assistants the questions your buyers ask and record whether you appear. It does not scale and it is the only method that shows you exactly what a buyer sees.

The measurement trap to avoid

Do not judge GEO content on sessions. That is the whole point of the shift, and it is the most common way good work gets killed. A definitional page earning citations and no clicks is performing. Reported as traffic, it looks like a failure.

The corollary matters too: do not judge commercial pages on citations. Pricing, comparison and alternatives pages still earn clicks and still convert, and optimizing them for quotability instead of conversion would be a straightforward mistake. Two content types, two scoreboards.

A realistic programme

If you are starting from nothing, in order.

  1. Fix retrieval first. If pages are not indexed, nothing else matters. GEO with broken technical SEO is decoration.
  2. Split your content in two and report each half against its own metric. Do this before writing anything new, or you will not be able to tell whether it worked.
  3. Restructure your highest-impression, lowest-click pages. They already rank. Making them extractable is cheaper than making new pages rank.
  4. Publish one original-data piece per quarter. Higher effort, highest return, and it compounds because other people cite it.
  5. Build definitional coverage of your category's vocabulary. Low competition, and these are the pages AI answers reach for.
  6. Pursue third-party mentions, particularly listicles and roundups in your category. When we checked which sources get cited for buyer-intent prompts in our space, they were overwhelmingly third-party roundups rather than vendor sites.
  7. Measure monthly, not weekly. Citation data is noisy and the lag between publishing and being cited runs to weeks.

How this changes the content calendar

The planning consequence is less obvious than the tactical one, and it is where most of the value is.

Stop planning by keyword volume alone

Volume tells you how many people search a term. It no longer tells you how many will arrive. A 10,000-volume definitional term in an AI-answered category may send a fraction of the visits a 500-volume commercial term does, while being far more valuable for presence.

Plan the two halves against different criteria. Commercial terms on volume and conversion potential. Definitional terms on how central they are to how buyers describe the problem, because that is what determines whether you get reached for.

Cover the vocabulary, not just the questions

Buyers and models both work from a category's vocabulary. If your space has thirty terms of art and you have written about six of them, you are absent from most of the conversations. Enumerating that vocabulary and covering it systematically is duller than chasing a big head term and reliably produces more.

The competitive check is straightforward: list the terms a competitor ranks for that you do not, filter to the definitional ones, and you have a backlog. We ran exactly this and found several at low difficulty where nothing on our site ranked at all.

Refresh on a schedule, not on a hunch

Recency is weighed for factual claims. Pricing, capability and market figures need real maintenance, and a page carrying a two-year-old number is a liability once it starts being quoted. Pick the pages that make dated claims and put them on an actual calendar.

Expect the reporting conversation to be hard

Content teams are measured on sessions. Telling a stakeholder that a page succeeded while sending no traffic is a difficult conversation, and it goes better with the query-length data in front of them than with an argument about industry trends. Showing that position improved while clicks collapsed, on their own property, moves people. Assertions do not.

Budget for the lag

Publishing to being cited runs to weeks, sometimes longer. A programme judged at thirty days will look like a failure regardless of quality. Ninety days is a fairer first read, and the honest framing at kickoff is that the first month produces no measurable result.

What GEO cannot do

Being honest about the ceiling, because the space is full of overclaiming.

It will not replace the traffic. If AI answers are absorbing clicks in your category, citations do not restore the sessions. They preserve presence at the research moment, which is worth something, but it is not the same asset.

The attribution is genuinely unsolved. A buyer who reads your name inside an AI answer and arrives three weeks later via a branded search is invisible to your analytics. Anyone quoting a confident ROI figure for GEO is estimating.

The surfaces move constantly. Tactics that work now may not in six months. Durable practice, being clear, specific, original and current, is a better bet than anything tuned to a particular model's current behaviour.

It does not fix a weak product story. Being cited accurately as a minor option in a category is not a win. Positioning is upstream of all of this.

Where Coworker AI fits

We run this on our own site, which is why this guide has our numbers in it rather than someone else's.

Coworker AI connects to 50+ tools, maintains organizational memory across them, and runs agents that act on what they find. It exposes that context over MCP so it is available inside the AI tools your team already uses. Pro is $29.99 per user per month, Max is $149.99, and Enterprise pricing is on request.

Book a demo if you want to talk through the content operation side of this.

Frequently asked questions

What is generative engine optimization?

The practice of making content likely to be selected, quoted and attributed inside AI-generated answers rather than clicked from a list of results. The term originates from a 2023 research paper that tested which content changes increased a source's visibility in generated responses.

Is GEO different from SEO?

The overlap is larger than the difference. Both need crawlable, well-structured, authoritative content. What differs is the objective, being cited rather than clicked, the unit, passages rather than pages, and the measurement. It is not a reason to run a separate content operation.

What is the difference between GEO and AEO?

In practice, very little. Answer engine optimization is usually used for the same objective, sometimes with more emphasis on direct-answer formats. The meaningful distinction is between optimizing for clicks and optimizing for citations, not between the acronyms.

How do I know if AI search is citing my site?

Imperfectly. Citation tracking tools sample prompt sets and disagree with each other. Search Console impressions act as a proxy for Google's surfaces. Prompting the assistants directly with your buyers' questions is crude but shows exactly what a buyer sees. Read trends rather than absolute numbers.

It is a low-cost, sensible thing to publish and a proposed convention rather than an adopted standard. Adoption by major providers is unconfirmed. Treat it as good hygiene, not a ranking lever, and be sceptical of anyone selling it as one.

How long does GEO take to work?

Longer than SEO reporting cycles suggest. The lag between publishing and appearing in AI answers runs to weeks, and citation data is noisy enough that a single week's movement means little. Measure monthly.

Should every page be optimized for citations?

No, and this is the most common mistake. Commercial-intent pages such as pricing, alternatives and comparisons still earn clicks and still convert, and should be optimized for conversion. Definitional and conversational content is where citation optimization pays. Applying one approach to both damages the half you get wrong.

Ready to get started?

Put Coworker to work inside your actual stack

Connect Salesforce, Slack, Jira, whatever you use, and run your first agent in minutes.