# Generative engine optimization: get read and cited by AI search

> Generative engine optimization (GEO) is the work of getting your pages read and cited by AI search like ChatGPT, Perplexity and Google AI Overviews. What is documented, what is opinion, what to do.

Updated 2026-09-26 · Technical SEO · HTML version: https://getreport.app/guides/generative-engine-optimization

Generative engine optimization (GEO) is the work of making your site something AI search can read, trust and cite: ChatGPT search, Perplexity, Claude, Microsoft Copilot and Google's AI Overviews and AI Mode. In practice it is mostly SEO with a few extra switches: letting the right AI crawlers in, serving your text in plain HTML, and writing pages that answer questions clearly enough to be quoted. This guide is for site owners and marketers who keep being told they need a GEO strategy. It separates what the AI operators actually document from what is still opinion, and ends with the steps you can take and check today.

## Quick answer

- **GEO is not a separate discipline from SEO.** Google says optimising for its generative AI features is "still SEO", and most AI search products retrieve pages through a search index before they write an answer.
- **Access comes first.** Each operator publishes its crawler names: OpenAI's OAI-SearchBot, Anthropic's Claude-SearchBot, PerplexityBot. If robots.txt or your firewall blocks them, you cannot be cited in those products.
- **Training and search are separate switches.** You can block GPTBot, ClaudeBot and Google-Extended (training) and still allow the search bots that cite and link you.
- **Serve the main text in the HTML.** The major AI crawlers do not run JavaScript, so content that appears only after scripts is invisible to them.
- **Write pages that answer questions directly**, with facts, numbers and sources in plain sentences. That is advice, not a documented ranking rule, but it is what AI answers quote.
- **llms.txt is optional.** It is a proposal, not a standard; Google Search ignores it, and no major AI search engine has committed to using it.
- **Check your site** with the free [AI crawler check](https://getreport.app/tools/ai-crawler-check), which shows what your robots.txt tells 18 AI crawlers and what they can read.

## What is generative engine optimization?

A generative engine is a search product that writes an answer instead of, or on top of, a list of links. You ask "which leather bag makers in Split make to order?", and the system searches, reads a handful of pages and writes a paragraph with citations. Generative engine optimization is everything that makes your page one of those it reads and cites.

The term comes from a 2023 research paper, *GEO: Generative Engine Optimization*, by Pranjal Aggarwal and colleagues from Princeton University and other institutions, presented at KDD 2024. They built a benchmark of queries, ran them through a generative engine they controlled, and tested rewriting source pages in different ways. Adding citations, quotations and statistics to a page raised its visibility in the generated answers by up to 40% in their setup. It is a useful result, but a lab one: it was measured on the authors' own engine and benchmark, not on ChatGPT, Perplexity or Google, which change their systems constantly.

## GEO, AEO and SEO: what the labels mean

You will meet several names for the same idea: GEO, AI search optimization, LLM SEO, LLM optimization, "AI SEO" and AEO, answer engine optimization. AEO is the oldest; it grew out of featured snippets and voice assistants, where the goal was to be the single answer lifted from a page. GEO came later with the research paper above and focuses on being cited in answers a language model writes.

None of these terms has an official definition or a specification. Google's guide to its generative AI features names AEO and GEO and says that, from Google Search's perspective, the work is still SEO. What genuinely differs is technical and small: separate AI crawlers, crawlers that do not run JavaScript, snippet controls that decide what can be quoted, and measuring citations instead of positions. Our comparison of [answer engine optimization vs GEO vs SEO](https://getreport.app/guides/answer-engine-optimization-vs-geo) goes through where each term comes from, what the operators document, and those four differences one by one.

## How AI search engines find and use your pages

Most AI search products work in two steps, and knowing them tells you what you can influence.

1. **Retrieval.** The system turns the question into one or more searches, runs them against a search index, and fetches the top pages. Google calls this retrieval-augmented generation and describes a "query fan-out": the system runs several related searches at once to cover different sides of the question. ChatGPT search, Perplexity and Claude run their own crawlers and search systems; Microsoft Copilot grounds its answers in Bing's index.
2. **Generation.** A language model reads the retrieved pages and writes the answer, citing some of them.

A third path is the model's **training data**: what it learned from pages crawled months or years earlier. That is how an assistant can talk about your brand without searching, and it is the part you influence least. You can keep new content out of it, as below, but you cannot edit what a model already learned.

So the levers are the same as in search, plus one. To be cited, a page has to be **allowed** (the crawler may fetch it), **readable** (the text is there without JavaScript), **retrievable** (it ranks in the index the product searches) and **quotable** (it states the answer clearly enough to be used).

## What is documented and what is opinion

GEO advice is full of confident claims with nothing behind them. This table separates what the operators themselves publish from what practitioners believe.

| Claim | Status | Source |
| --- | --- | --- |
| AI Overviews and AI Mode have no extra requirements beyond normal Search indexing and snippet eligibility | Documented | Google Search Central, "AI features and your website" |
| Google Search ignores llms.txt and needs no special AI files or markup | Documented | Google's guide to optimizing for generative AI features in Search, May 2026 |
| Structured data is not required for Google's AI features | Documented | Same Google guide |
| Blocking OAI-SearchBot keeps a site out of ChatGPT search answers | Documented | OpenAI's crawler documentation |
| Google-Extended controls Gemini training and grounding, not Google Search | Documented | Google's list of common crawlers |
| Schema markup helps Microsoft's language models understand content | Stated by Microsoft | Fabrice Canel (Bing) at SMX Munich, March 2025 |
| Adding statistics, quotations and citations raises visibility in generated answers | Research result, one lab setup | Aggarwal et al., KDD 2024 |
| Mentions of your brand on other trusted sites help AI answers name you | Widely believed; no operator documents it | Practitioner opinion |
| Content split into short, self-contained sections is quoted more often | Opinion; Google says there is no need to chunk content | Practitioner opinion vs Google's guide |
| llms.txt helps you get cited | Unproven; no major AI search engine has said it reads the file | llmstxt.org proposal |

Our view, labelled as such: the documented items are enough to act on, and they are the cheap ones. Anything sold as a GEO "hack" beyond them deserves a question about its evidence.

## Step 1: let the right AI crawlers in

Every major operator runs separate agents for training, for its search index and for fetching a page when a user asks. They are documented by the operators:

| Operator | Training | AI search | User-triggered fetch |
| --- | --- | --- | --- |
| OpenAI | GPTBot | OAI-SearchBot | ChatGPT-User |
| Anthropic | ClaudeBot | Claude-SearchBot | Claude-User |
| Perplexity | (none; PerplexityBot is not used for foundation models) | PerplexityBot | Perplexity-User |
| Google | Google-Extended (a robots.txt token, not a crawler) | Googlebot, as for Search | — |

What the operators say about robots.txt differs, and the details matter:

- **OpenAI** documents that sites that disallow OAI-SearchBot will not be shown in ChatGPT search answers, though they can still appear as navigational links, and that changes take about 24 hours to apply. For ChatGPT-User it says that, because a user starts the action, robots.txt rules may not apply.
- **Anthropic** says ClaudeBot, Claude-SearchBot and Claude-User all respect robots.txt, and that blocking ClaudeBot does not affect the other two.
- **Perplexity** says PerplexityBot respects robots.txt and that Perplexity-User, which fetches a page because a user asked, generally ignores it.
- **Google** says Google-Extended controls whether content may be used to train Gemini models and for grounding, and does not affect inclusion or ranking in Google Search. AI Overviews and AI Mode are part of Search: the controls there are Googlebot access, the usual `noindex`, `nosnippet`, `data-nosnippet` and `max-snippet` directives, and since 2026 the Search generative AI control in Search Console, which opts a whole site out.

The decision itself is a business one. A shop usually wants to be cited everywhere; a publisher may block training and allow search. Treating every AI bot as one thing is the most common mistake: a 2023-era robots.txt that blocks GPTBot does nothing about OAI-SearchBot, and a `Disallow: /` meant for AI bots under `User-agent: *` also blocks Googlebot. [AI crawlers and robots.txt](https://getreport.app/guides/ai-crawlers-and-robots-txt) has complete robots.txt files for four kinds of site, the rules of how groups are read, and how to enforce a block at the edge.

Check the firewall as well. Cloudflare and some security plugins can block AI crawlers regardless of robots.txt, so a site can say "allowed" and still never be fetched.

## Step 2: make your content readable without JavaScript

An analysis of AI crawler traffic published by Vercel and MERJ in December 2024 found that none of the major AI crawlers rendered JavaScript, including OpenAI's, Anthropic's, Meta's, ByteDance's and Perplexity's. Some fetched script files, but read them as text rather than running them. Google's Gemini is the exception, because it uses Googlebot's rendering.

For a site built as a client-rendered app, where the HTML the server sends is an empty `<div id="root">` and scripts fill it in, that means most AI crawlers see nothing to quote. Server-side rendering or static generation fixes it: the text, headings, links and meta tags are in the HTML from the start, and scripts add interactivity on top. The same fix helps Google index faster and makes link previews work. [JavaScript-rendered content and Google](https://getreport.app/guides/javascript-rendered-content-and-google) covers the rendering options for Google.

The gaps are often smaller than a full client-rendered app: prices fetched from an API after load, reviews injected by a widget, JSON-LD added through a tag manager, or a canonical changed by a script. Each is in the rendered page but missing from the HTML an AI crawler reads. Browser agents that act for a user can run scripts, but they do not build the indexes answers come from. Our guide to [whether AI crawlers can execute JavaScript](https://getreport.app/guides/can-ai-crawlers-render-javascript) has the crawler-by-crawler table, the tests you can run with View Source or curl, and the fix for each common framework.

Menus, carousels and widgets that need JavaScript are fine. Product descriptions, prices, article text and answers are not.

## Step 3: be findable in the indexes AI search uses

Retrieval runs on search indexes, so classic technical SEO is the foundation of GEO:

- **Indexable pages**: no stray `noindex`, correct canonicals, pages in the XML sitemap and the sitemap listed in robots.txt.
- **Google Search** feeds AI Overviews and AI Mode. Google says a page must be indexed and eligible to show a snippet to appear there, and nothing more.
- **Bing** feeds Microsoft Copilot, and Bing's results are widely reported to be one of the sources behind ChatGPT search, alongside OpenAI's own crawling. Verify the site in Bing Webmaster Tools and submit the sitemap; many sites never have.
- **Speed and reliability** matter for crawlers too: a page that times out or answers 5xx is not fetched.

### ChatGPT search

ChatGPT is the AI product most people mean when they ask about AI visibility, and it works differently from Google's features. When a question needs current information it searches the web through third-party search providers, with Bing as its announced partner since 2023, alongside OpenAI's own crawler, OAI-SearchBot. OpenAI documents that blocking OAI-SearchBot keeps a site out of ChatGPT search answers; GPTBot, the training crawler, is a separate switch.

In practice that means four jobs: allow OAI-SearchBot in robots.txt and at the firewall, get indexed in Bing as well as Google, serve the text without JavaScript, and measure with the `utm_source=chatgpt.com` parameter ChatGPT adds to its links. The guide to [SEO for ChatGPT](https://getreport.app/guides/how-to-rank-in-chatgpt) goes through each step, plus product feeds for shops and what does not help.

### Google AI Overviews and AI Mode

Google's AI features are the largest AI search surface, and the one with the clearest rules. They are built from Google's own index: the system retrieves pages with its core ranking systems, often running several related searches for one question ("query fan-out"), and links some of them as sources. Google's May 2026 guide says there is nothing extra to add, no schema, no llms.txt, no special writing style, and warns that mass-produced pages for every phrasing of a question count as scaled content abuse.

What you do control is whether you take part. Since mid-2026, the Search generative AI control in Search Console removes a whole site from AI Overviews and AI Mode without affecting normal rankings, and `nosnippet`, `data-nosnippet` and `max-snippet` limit what can be quoted. Google-Extended does not affect either. Our guide on [how to rank in AI Overviews](https://getreport.app/guides/google-ai-overviews-seo) covers how pages are picked, the eligibility checks, those controls and what makes a page more likely to be linked.

## Step 4: write pages worth quoting

This is where most GEO advice lives and where evidence is thinnest. Google's own guide is blunt: create content that is "unique, compelling, and useful", and it warns against commodity pages such as generic tip lists that repeat what everyone else says. It also says you do not need to write in a special way for AI or chop content into small pieces.

Within that, a few habits help any reader, human or model, and they match what the GEO research tested:

- **Answer the question in the first sentence under the heading**, then explain. "Delivery to Germany takes 3–5 working days and costs 9 EUR" is quotable; "We pride ourselves on fast shipping" is not.
- **State facts with numbers and units**: prices, dimensions, opening hours, limits, dates.
- **Cite your sources** by name when you use a number or a claim that is not yours.
- **Say who wrote it and when**, with an author, a date and a way to contact you. It helps people judge the page, and AI answers often mention the source.
- **Cover what is unique to you**: your prices, your process, your data, your experience. That is the one thing an AI answer cannot get from someone else.

Structured data belongs here too. It states facts such as price, availability and opening hours in a form machines read directly. Google says it is not required for its AI features; Microsoft says it helps its models. It is worth having for search anyway, and [schema markup](https://getreport.app/guides/schema-markup) covers which types to add.

The record on schema and AI is short. Google says structured data is not required for AI Overviews or AI Mode, Microsoft says schema helps its models understand content, and OpenAI says ChatGPT's shopping results use structured metadata such as price and description. Anthropic and Perplexity have said nothing. Nobody documents schema as a citation factor, and there is no special "AI schema" type, whatever some plugins suggest.

What matters in practice is placement and accuracy: JSON-LD in the server's HTML rather than injected by a tag manager, facts that match the visible page, and one consistent entity per page. The guide to [schema markup for AI search](https://getreport.app/guides/schema-markup-for-ai-search) lays out each operator's statement with its source, the types worth adding and how to check them.

Being mentioned elsewhere, in reviews, press, directories and forums, is widely believed to shape which brands AI answers name. No operator documents that, so treat it as opinion; Google explicitly warns against manufacturing inauthentic mentions.

## Step 5: add llms.txt, if you like, as a cheap bet

llms.txt is a Markdown file at `/llms.txt` that lists your most useful pages with a line on what each answers. It was proposed at llmstxt.org in September 2024. It is not a web standard, and the situation in September 2026 is clear on one side: Google's guide says Google Search ignores it. No other major AI search engine has said it reads the file for search or citations. Coding assistants and documentation tools do use it, which is why it is common on developer documentation sites.

It takes about 15 minutes, so it is a reasonable bet with a small cost, not a GEO strategy. [llms.txt: what it is and how to write one](https://getreport.app/guides/llms-txt-what-it-is-and-how-to-write-one) covers the format, examples and placement, and the generator drafts one from your sitemap:

> **Free tool:** [llms.txt generator: built from your sitemap](https://getreport.app/tools/llms-txt-generator): Free llms.txt generator: we read your sitemap, draft the file with real page titles in sections, and you edit it in your browser with live checks.

It reads your robots.txt and sitemap, groups up to 500 pages into sections with their real titles, and lets you edit the draft in your browser, checked against the llmstxt.org format as you type.

## How to check your site's AI readiness

> **Free tool:** [Check which AI crawlers can read your site](https://getreport.app/tools/ai-crawler-check): See what your robots.txt tells 18 AI crawlers, such as GPTBot, ClaudeBot and PerplexityBot, whether your llms.txt is valid and what needs JavaScript. Free.

The check reads your robots.txt and, for each of 18 AI user agents from OpenAI, Anthropic, Google, Perplexity, Apple, Meta, Amazon, ByteDance, Common Crawl and others, answers the question the crawler asks: may I fetch this page? Crawlers are grouped by purpose (training, AI search, user fetches), so you can see at a glance whether you are blocking the ones that cite you. It also validates `/llms.txt`, looks for AI opt-out tags, and compares the raw HTML with the rendered page to show how much text needs JavaScript. The same findings appear in every getReport report:

> **Check: AI crawler access.** AI assistants (ChatGPT, Claude, Perplexity) fetch pages with their own user agents. Google's AI Overviews use normal Googlebot, so only blocking Googlebot removes you from them, and that removes you from Search too. Blocking them keeps your content out of AI answers; allowing them can bring citations and visitors. Either is a valid choice, as long as it is the one you meant.
>
> 1. Decide per crawler: training bots (GPTBot, ClaudeBot, Google-Extended, CCBot) feed models; search bots (OAI-SearchBot, PerplexityBot) feed answers that link back; user bots (ChatGPT-User, Claude-User) fetch a page because someone asked. Google-Extended is a robots.txt token, not a crawler: it controls use in Gemini training, not AI Overviews.
> 2. Edit robots.txt: a "User-agent: GPTBot" group with "Disallow: /" blocks it; leave it out of robots.txt (or "Allow: /") to permit it. Rules for "*" apply to any crawler without its own group.

> **Check: The visible text is present without JavaScript.** Google renders JavaScript later and with a budget, so text that only appears after scripts run can be indexed late or not at all. Other search engines and link previews may never see it.
>
> 1. Serve the main content in the HTML (server-side rendering or static generation) and use JavaScript only to enhance it.
> 2. Check the difference in the technical detail; menus and widgets are fine, headlines and body copy are not.

> **Check: llms.txt found.** /llms.txt is a plain-text index of your most useful pages written for AI assistants, in the format proposed at llmstxt.org. It is optional and new: a few tools read it, but no major assistant has confirmed using it and Google has said Search does not. It costs about ten minutes, so treat it as a cheap bet, not a ranking factor.
>
> 1. Create /llms.txt: an H1 with the site name, a one-line blockquote summary, then "## Section" headings with Markdown links to your key pages and a short description each.
> 2. Serve it as text/plain or text/markdown; optionally add /llms-full.txt with the full text of those pages.

> **Check: No AI opt-out directive on the page.** A "noai" or "noimageai" robots directive or a TDM reservation tag asks AI systems not to use the page for training. Support varies by operator; it is a signal, not a lock, but it documents your choice.
>
> 1. Keep it if it is deliberate. To opt back in, remove the noai/noimageai tokens from the robots meta tag or X-Robots-Tag header and the tdm-reservation meta tag.
> 2. For a real block, pair it with robots.txt rules for the crawlers you want out.

What the check cannot do is tell you what ChatGPT or Perplexity actually say about you. It shows what your site tells each crawler, not how the crawler behaves or which answers cite you. For who really visits, read your server logs: [server log analysis for Googlebot and AI bots](https://getreport.app/guides/server-log-analysis-googlebot-and-ai-bots) shows how to count requests from each AI crawler, which pages they fetch and the status codes they get.

## Measuring visibility in AI search

Measurement is the weakest part of GEO, but it improved in 2026. Google added a Generative AI performance report to Search Console, rolled out to all sites by 31 August 2026, which shows impressions in AI Overviews and AI Mode by page, country, device and date. Bing Webmaster Tools' AI Performance report goes further for Microsoft's products: it counts how often Copilot cited each page and lists the "grounding queries" it ran to find them. Neither shows clicks per answer, and ChatGPT, Claude and Perplexity offer no publisher report at all.

For everything else you piece the picture together: AI referrals in analytics (ChatGPT tags its links with `utm_source=chatgpt.com`, though some apps send no referrer), user-triggered fetches such as ChatGPT-User in your server log, and prompt-tracking tools that sample answers to a fixed list of questions. Those samples vary between runs and users, so they show trends, not positions.

The guide to [AI search visibility](https://getreport.app/guides/ai-search-visibility) sets out each source with its limits, a one-hour baseline you can repeat monthly, and the steps that improve how often your brand is cited, ordered by how well they are documented.

## Generative engine optimization vs traditional SEO

| | Traditional SEO | GEO |
| --- | --- | --- |
| Goal | Rank a page in a list of links | Be read and cited inside a written answer |
| Where it happens | Google, Bing and other search engines | AI Overviews, AI Mode, ChatGPT, Perplexity, Claude, Copilot |
| Crawlers | Googlebot, Bingbot | The same, plus OAI-SearchBot, Claude-SearchBot, PerplexityBot and others |
| JavaScript | Google renders it, with a delay | Most AI crawlers do not render it |
| Measured by | Rankings, impressions, clicks | Citations and mentions; impressions in Search Console; referrals |
| Foundation | Crawlable, indexable, relevant, useful pages | The same |

The last row is the important one. GEO adds crawler decisions, rendering and a different way of measuring, but the page that gets cited is almost always a page that would deserve to rank.

## Common mistakes

- **Blocking AI search bots while trying to be cited.** A blanket AI block in robots.txt or at Cloudflare takes you out of the answers you want to appear in.
- **Blocking Google-Extended to leave AI Overviews.** It does not do that; it only controls Gemini training and grounding. AI Overviews follow your Search settings.
- **A client-rendered site with no server HTML.** Most AI crawlers see an empty page.
- **Buying a GEO audit that is really a rank tracker.** Ask what it measures and how often the same question gives a different answer.
- **Rewriting every page "for AI".** Google says there is no need; make the page clearer for people and it is clearer for models too.
- **Treating llms.txt as the strategy.** It is a 15-minute extra, not a substitute for access, readable HTML and good pages.

## Questions people ask

### What is generative engine optimization?

Generative engine optimization (GEO) is the work of getting your pages read and cited by AI search products such as ChatGPT search, Perplexity, Claude, Copilot and Google's AI Overviews. It means letting their crawlers in, serving text in plain HTML, ranking in the indexes they search, and writing pages that answer questions clearly. The term comes from a 2023 research paper.

### What is AI SEO?

AI SEO is a loose label with two meanings: optimising a site so AI search products read and cite it, which is the same as generative engine optimization, or using AI tools to do SEO work, such as drafting content or clustering keywords. When someone offers AI SEO, ask which one they mean, and what evidence they have that it changes citations or traffic.

### How is generative engine optimization different from traditional SEO?

GEO aims to be cited inside a written answer rather than to rank in a list of links. It adds a few things: deciding which AI crawlers to allow, serving content without JavaScript because most AI crawlers do not run it, and measuring citations instead of positions. The foundation is the same, and Google describes optimising for its AI features as still SEO.

### Is generative engine optimization worth it for a small business?

Yes, the documented part is cheap: allow AI search crawlers, keep your main text in the HTML, be indexed in Google and Bing, and answer the questions customers ask on your own pages. That takes hours, not a retainer. Paying for ongoing GEO services is harder to justify for a small site, because most of what they sell beyond those basics has no published evidence.

### Does blocking GPTBot remove my site from ChatGPT?

No. GPTBot collects pages for training OpenAI's models. ChatGPT search uses a different crawler, OAI-SearchBot, and OpenAI says sites that block OAI-SearchBot are not shown in ChatGPT search answers. You can block GPTBot and allow OAI-SearchBot in robots.txt to stay out of training and still be cited. Changes take about a day to apply.

### How can I see whether AI search mentions my site?

Partly. Search Console's Generative AI performance report shows your impressions in Google's AI Overviews and AI Mode by page. Analytics shows visits referred by ChatGPT, Perplexity, Claude and Copilot, although some arrive without a referrer. For other products, prompt-tracking tools sample answers to a fixed list of questions; results vary between runs, so watch trends, not single answers.

### Is GEO just a new name for SEO?

Mostly, and Google says as much: from its perspective, optimising for generative AI search is still SEO. The parts that are genuinely new are small but real: separate crawlers for training and AI search with their own robots.txt names, crawlers that do not run JavaScript, and answers that cite rather than rank. Those deserve attention; a whole new discipline does not, in our view.
