What is meta-externalagent?
meta-externalagent is Meta's crawler for gathering data to train its AI models and power features like Meta AI, consolidating jobs that used to run under several separate, narrower Meta bots.
Meta consolidated its AI-related crawling under one user agent, meta-externalagent, rather than running separate bots for different purposes the way OpenAI splits GPTBot from OAI-SearchBot. That means a robots.txt rule targeting meta-externalagent is currently an all-or-nothing decision for Meta's AI uses of a site's content, there's no equivalent line yet to allow retrieval only while opting a page out of training.
This bot is distinct from the older Facebook crawler that fetches pages when a link gets shared on Facebook or Instagram, which still runs under its own user agent and exists to build link previews, not to feed AI models. Blocking meta-externalagent doesn't change how a page's preview card looks when someone pastes its URL into a Facebook post.
Related
- GPTBotGPTBot is OpenAI's web crawler. It reads pages to train and ground OpenAI's models, including ChatGPT.
- Training vs search crawlersA training crawler collects data models learn from. A search crawler builds the index an assistant cites when answering. Blocking one has nothing to do with the other.
- AI crawlerAn AI crawler is a bot operated by an AI company to fetch web pages for training, indexing, or answering a live question.
- User agentA user agent is the name a client sends to identify itself when requesting a page. robots.txt rules are written against these names, which is why blocking AI crawlers means naming each one.
Want to know where you actually stand on this? Run a free visibility check or try the free tools.