To get cited by ChatGPT, allow at least OAI-SearchBot and ChatGPT-User in your robots.txt, along with Bingbot. GPTBot is used to train models: blocking it does not stop you being cited in searches, but makes it less likely that ChatGPT knows your brand.
This article complements our GEO guide.
Three robots, three roles
ChatGPT uses several robots, each with a distinct role. Mixing them up causes most configuration mistakes.
| Robot | Role | Effect of blocking it |
|---|---|---|
| OAI-SearchBot | Crawls the web for ChatGPT search | Your pages can no longer appear as sources in answers with search |
| ChatGPT-User | Reads a page on demand, when a user or ChatGPT wants to open it during a conversation | ChatGPT cannot open your page when someone asks for it |
| GPTBot | Collects pages to train models | Your content no longer feeds training; no direct effect on citations in search |
The key point: OAI-SearchBot and ChatGPT-User decide citations, GPTBot decides what the model knows in general.
Should you block GPTBot?
It is an editorial decision. Some news publishers block it so their articles do not feed training without compensation. For a small business that wants to be known and recommended, the reasoning is often the opposite: the more your business appears in training data, the more likely the model knows you when it answers without searching.
Our recommendation for a commercial site: allow all three. If you prefer to block GPTBot, do it explicitly and keep OAI-SearchBot and ChatGPT-User open.
What about other AI robots?
| Robot | Service | Advice |
|---|---|---|
| Bingbot | Bing’s index, which ChatGPT relies on heavily | Allow, without fail |
| PerplexityBot | Perplexity search | Allow to be cited there |
| Claude-SearchBot | Claude search | Allow to be cited there |
| ClaudeBot | Training of Claude | Same logic as GPTBot |
| Google-Extended | Use of your pages by Google’s AI (Gemini) | No effect on your Google rankings |
Watch out for Bingbot: many analyses show that ChatGPT’s citations largely overlap Bing’s top results (Stackmatix, Yoast). Blocking Bingbot cuts you off from a large share of citations.
What does a good robots.txt look like?
A robots.txt open to AI robots and search engines:
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: GPTBot
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: *
Allow: /
Disallow: /wp-admin/
Sitemap: https://mysite.com/sitemap.xml
And if you only want to block training:
User-agent: GPTBot
Disallow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
A detail that often trips people up: a robot follows the group that names it most precisely. If you have a User-agent: GPTBot group with Disallow: /, the User-agent: * rule no longer applies to GPTBot. Conversely, a User-agent: * group with Disallow: / blocks every robot not named elsewhere, including OAI-SearchBot.
What are the most common blocking mistakes?
- A
Disallow: /left over from a site under construction, never removed at launch. - A WordPress security plugin that adds rules against “bad bots” and lumps AI robots in with them.
- The host’s or CDN’s firewall blocking AI robots without going through robots.txt. Your file is perfect, but robots get a 403 error. Check your host’s bot protection settings.
- A robots.txt that cannot be found on the www or non-www version, or over HTTP.
- Content shown only with JavaScript. The robot gets in but sees nothing: many AI robots do not run JavaScript.
How to check in 30 seconds?
Enter your address in our robots.txt checker: it reads your file and tells you, robot by robot, which are allowed and which are blocked, with the lines to add. To go further (firewall, JavaScript, tags, speed), run your site’s free audit.
Once robots are unblocked, it takes a few days to a few weeks for them to come back. ChatGPT tracking then shows you, week after week, whether your pages start appearing in answers.