Skip to content
Guides

Does publishing on Substack or Medium hurt your AI recommendations?

6 min read

You published a detailed technical guide on your brand website last week. Yesterday, you asked ChatGPT, Perplexity, and Google AI Overviews a prompt your guide answers directly. The engines did not cite your site. Instead, they cited a two-year-old post on Medium and a newsletter hosted on Substack.

Now you are wondering whether to move your technical content to Substack or Medium to get recommended by AI search engines.

The short answer is that publishing on Medium or Substack often improves your immediate visibility in live AI search results, but it limits your long-term brand authority and exposes you to platform risks you cannot fix. Understanding why requires looking at how retrieval-augmented generation works under the hood.

The immediate trade-off: domain authority versus domain control

AI models do not search the entire web from scratch every time a user types a prompt. Engines like Perplexity, ChatGPT with Search, and Google AI Overviews rely on underlying search indices and live retrieval bots to fetch pages.

When an AI engine executes a search query to answer a user prompt, it prioritizes domains with high crawl frequency and established authority. Platforms like Substack and Medium have millions of inbound links and high crawl priority. When you post an article on medium.com or substack.com, search crawlers index it within minutes or hours.

If you publish the same article on a new standalone domain with low domain authority, search indexers might take days or weeks to index the page. If an AI engine uses a search API that has not indexed your new URL yet, the engine cannot retrieve your page as a source.

In the short term, publishing on Medium or Substack gives your content a higher probability of being retrieved for broad industry prompts. But that visibility comes with structural trade-offs.

Training crawlers versus answer crawlers

To evaluate platform risks, you must understand the technical difference between training crawlers and answer crawlers. Confusing these two is the most common error made when optimizing for AI recommendations.

Training crawlers gather web data to build future model weights. Examples include GPTBot, ClaudeBot, and Google-Extended. Blocking a training crawler stops an AI company from using your content to train its next foundation model, but it does not stop that engine from citing your site in live search today.

Answer crawlers fetch web pages in real time to answer a specific user query right now. Examples include ChatGPT-User, OAI-SearchBot, Perplexity-User, PerplexityBot, Claude-SearchBot, and Claude-User. Blocking an answer crawler prevents the engine from reading your content during a live search run, eliminating your citations immediately.

When you publish on your own domain, you control your robots.txt file and can decide which crawlers to allow or block. When you publish on Substack or Medium, you inherit their platform-wide robots.txt configuration.

In our dataset scan on 2026-08-08 across a panel of 54,082 domains, 33,670 domains returned a readable robots.txt file. Of those readable domains, 5,497 blocked at least one AI crawler, and 2,576 blocked at least one answer crawler.

Looking at answer crawlers specifically across those 33,670 readable domains:

  • 2,147 blocked PerplexityBot
  • 2,074 blocked ChatGPT-User
  • 1,579 blocked OAI-SearchBot
  • 1,402 blocked Perplexity-User
  • 1,400 blocked Claude-SearchBot
  • 1,384 blocked Claude-User
  • 610 blocked Googlebot

Compare those figures to training crawlers across the same 33,670 readable domains:

  • 5,080 blocked GPTBot
  • 4,603 blocked ClaudeBot
  • 4,275 blocked Google-Extended

A separate census of Y Combinator companies conducted on 2026-08-07 evaluated 4,226 active company domains. Of the 3,755 domains with readable configuration files, 253 blocked at least one AI crawler.

If Medium or Substack alters its global robots.txt file to block answer crawlers like ChatGPT-User or OAI-SearchBot, every article hosted on their platform loses live AI citations across those engines overnight. As a tenant on their domain, you have zero control over that decision.

Attribution loss and entity resolution

Even when an AI engine retrieves your content from Substack or Medium, you face an attribution problem.

AI models use entity resolution to connect facts to specific brands. When an answer crawler reads a page on yourbrand.com, the model connects the concepts on that page directly to your corporate domain and brand entity.

When an answer crawler reads an article on medium.com/@yourbrand/article, the language model often credits Medium as the authoritative publisher rather than your brand. In live runs across Perplexity and Google AI Overviews, citations frequently display as "Medium" or "Substack" in the UI footnotes rather than naming your company.

If a user asks an AI engine for a product recommendation, the model synthesizes sources to evaluate vendor credibility. A source published on an official company domain carries direct brand weight. A source published on a third-party blogging platform is often parsed as secondary commentary or independent opinion, which can dilute how strongly the engine associates your brand with the solution.

Using a custom domain on Substack (such as newsletter.yourbrand.com) mitigates some attribution loss by retaining your domain name in the cited URL. However, you remain bound by Substack's technical infrastructure and crawler access policies.

Testing content location performance

Because AI engines are non-deterministic, you should test how your specific audience prompts behave rather than relying on general assumptions.

To test whether hosting content on your own domain outperforms Substack for your specific terms:

  1. Draft two technical, highly specific articles addressing queries your prospective buyers ask.
  2. Publish one article on your primary domain and the other on a platform like Substack or Medium.
  3. Allow several days for indexing across standard search engines.
  4. Run live prompt tests across ChatGPT, Perplexity, Gemini, Claude, and Google AI Overviews to observe citation rates.

Do not evaluate performance by asking a prompt once. AI engines vary their search queries and retrieval sources from run to run. A single run might cite Substack, while the next run cites a standalone domain.

At Standing, we measure brand presence by running a fixed prompt basket five times per question across each supported engine. We calculate the output using a Wilson interval to account for variance. For example, a result might show a citation score of 34, plus or minus 6, out of 100 runs at a 95% confidence interval. Competitors typically run each prompt only once, which produces a single point estimate without a confidence band and masks underlying non-deterministic variance.

Our Track tier costs $100 per month for 3 domains. Our Optimize tier costs $300 per month for 5 domains with custom prompts. Our Agency tier costs $500 per month for 50 domains with a monthly re-scan.

What not to buy

As you evaluate where to publish, avoid vendors selling shortcuts.

Do not buy services promising guaranteed citations or fixed rankings in ChatGPT or Perplexity. No vendor can guarantee a citation because retrieval algorithms and model outputs change continuously.

Do not spend money on plugins or agencies claiming that adding schema markup or llms.txt files will force AI engines to recommend your business. We have found no reliable evidence that schema markup or llms.txt files increase AI citation rates, and technical files do not overcome a lack of domain authority or poor retrieval relevance.

If your primary site is brand new and receives zero search crawler traffic, publishing temporary educational content on Medium or Substack can establish initial visibility. But for long-term brand equity, entity recognition, and complete control over answer crawler access, publishing on your own domain remains the necessary foundation.

See where you actually stand

Run a free check on your domain. Five AI surfaces, the real buying questions, about a minute. No signup.

Run a free check

Keep reading