Is GEO a scam?
Someone on your board or executive team pasted a prompt into ChatGPT, saw a competitor named instead of you, and sent an email asking what your Generative Engine Optimization strategy is. Within a week, three vendors pitched you monthly retainers promising to rank your brand across AI search engines.
Most of what those agencies offer is snake oil. Generative Engine Optimization (GEO) as currently pitched by the market is largely a repackaged set of traditional SEO tactics mixed with promises that no vendor can keep. Yet under the hype, there is a narrow set of technical realities that govern whether an AI engine cites your business.
Understanding the difference saves you from spending thousands of dollars on useless deliverables while fixing the actual reasons models ignore you.
What is real in GEO
AI engines like ChatGPT, Perplexity, Gemini, Claude, and Google AI Overviews do not invent every answer from static memory. When asked about products or services, they perform real-time web searches or retrieve indexed source documents, synthesize those documents, and generate a response. Three things genuinely affect this process: crawler access, source site presence, and statistical measurement.
First, your site must be reachable by answer crawlers. If an engine cannot fetch your site when a user asks a question, it cannot cite your page in real time.
Many companies accidentally block these crawlers in their robots.txt files. Crucially, you must distinguish between training crawlers and answer crawlers. Training crawlers, such as GPTBot, ClaudeBot, and Google-Extended, scrape web content to build future base models. Blocking a training crawler stops an engine from training on your content for its next model release, but it does not stop the engine from visiting your site today to answer a query.
Answer crawlers, such as ChatGPT-User, Claude-SearchBot, OAI-SearchBot, and Perplexity-User, fetch live web pages during an active user conversation. If you block an answer crawler, the engine cannot read your site to generate a real-time citation.
In our crawler index run on 2026-08-08, out of a panel of 54,082 domains, 33,670 returned a readable robots.txt file. Of those 33,670 domains, 5,497 block at least one AI agent, and 2,576 block at least one answer crawler.
The breakdown of blocked agents out of those 33,670 readable domains shows how frequently companies confuse these roles:
- GPTBot (training): 5,080 domains block it
- ClaudeBot (training): 4,603 domains block it
- Google-Extended (training): 4275 domains block it
- PerplexityBot (search): 2,147 domains block it
- ChatGPT-User (search): 2,074 domains block it
- OAI-SearchBot (search): 1,579 domains block it
- Perplexity-User (search): 1,402 domains block it
- Claude-SearchBot (search): 1,400 domains block it
- Claude-User (search): 1,384 domains block it
- Googlebot (other): 610 domains block it
This confusion extends to early-stage technology companies. In our YC census run on 2026-08-07 covering 4,226 active Y Combinator companies, 3,755 had readable robots.txt files. Of those 3,755 domains, 253 block at least one AI crawler.
Checking and fixing your robots.txt to permit answer crawlers is concrete, immediate, and verifiable.
Second, your brand presence on third-party sources directly changes AI outputs. When Perplexity or ChatGPT answers a query like "best CRM for real estate," it queries search APIs, retrieves top-ranking articles, comparison sites, and forum threads, and summarizes them. If your brand is absent from the specific pages retrieved for that topic, the engine will not invent your presence. Earned media, inclusion on buyer guides, and active discussions on platforms engines reference are real levers that change model outputs.
Third, measurement is possible, but only if conducted with statistical rigor. Large language models are non-deterministic. Asking ChatGPT a single question once tells you almost nothing, because the model may give a different answer on the next request. Real measurement requires querying an engine multiple times across a fixed basket of prompts and calculating a confidence interval.
At Standing, our method uses a fixed prompt basket, running each question five times per engine across ChatGPT, Perplexity, Gemini, Claude, and Google AI Overviews, then calculating a Wilson interval on the result. A score must be expressed as a range, such as 34, plus or minus 6, to account for model variance. Competitors typically run each prompt once, reporting absolute numbers that give a false impression of stability.
What is fake in GEO
The market for GEO is flooded with vendors selling tactics that range from useless to impossible.
The most common pitch is fixed placements or assured positions in generative answers. No agency or software tool can deliver a specific placement or mention in an LLM. Because these engines synthesize responses dynamically using probabilistic sampling, outputs vary by location, query phrasing, time, and random seed. Any vendor offering fixed outcomes is selling a result they cannot deliver.
The second common delusion is that schema markup and custom files like llms.txt act as direct citation levers. Schema markup helps traditional search engines understand structured data for rich snippets, but LLMs synthesize unstructured text directly from web pages. There is no evidence that adding JSON-LD schema or an llms.txt file increases the likelihood of being cited by an AI engine. In fact, relying on hidden metadata while ignoring the body copy of your pages does nothing for citation rates.
The third indicator of snake oil is any performance report that gives a single static score without a confidence band. A reporting platform that tells you your brand visibility is a single number like 42 without stating the variance or interval is showing you noise, not data.
Why the dishonest version sells
The honest version of GEO is slow, technical, and hard to outsource entirely. It requires auditing crawler logs, removing explicit blocks in robots.txt, building real brand authority on high-quality third-party sites, writing detailed documentation, and measuring progress through statistical sampling over time. It offers no silver bullets and no simple code snippets that double your visibility overnight.
The dishonest version of GEO sells because it matches what buyers want to hear. Executives want a fast fix for a board member's prompt. Agencies meet that demand by offering low-effort deliverables like injecting schema markup, uploading an llms.txt file, or generating repetitive AI text aimed at keyword stuffing.
These tactics sell because they feel like traditional SEO deliverables from ten years ago. You can buy them on a monthly retainer, receive a monthly PDF with a single unsampled visibility score, and pretend the problem is solved. But because LLMs operate on live search index retrieval and probabilistic synthesis rather than keyword density and metadata tags, those deliverables do not move the needle.
If you choose to track your visibility, demand statistical transparency. At Standing, our plans start at $100 per month for Track (covering 3 domains), $300 per month for Optimize (5 domains with custom prompt baskets), and $500 per month for Agency (50 domains with a monthly re-scan). We publish our exact methodology and confidence bands because measuring probabilistic systems requires acknowledging uncertainty.
When evaluating a GEO proposal, ask three questions: Does this vendor claim deterministic outcomes? Do they treat schema or llms.txt as a ranking factor? And do their reports show a confidence interval? If they answer yes to the first two or no to the third, you are looking at a vendor selling quick fixes for a complex problem.
See where you actually stand
Run a free check on your domain. Five AI surfaces, the real buying questions, about a minute. No signup.
Run a free check