How ChatGPT chooses
What it reads, and when.
A customer asks ChatGPT who to call for a blocked drain in their suburb, and it names three businesses. Whether yours is one of them comes down to two very different things: what the model already learned before its cut-off, and what a live web search finds the moment the question is asked. OpenAI publishes more about the second than most people realise. Here is what its own documentation says, crawler by crawler, and the part a business can actually influence.
Every claim here about how ChatGPT and OpenAI behave is cited to OpenAI's own developer documentation, read on 25 September 2026 and re-read on 28 September 2026. Where OpenAI does not say how it chooses, neither do we. Updated 28 September 2026. 9-minute read. How we write these guides.
The short answer.
ChatGPT answers a question about businesses in one of two ways. It draws on what the model learned before its knowledge cut-off date[1], or it uses ChatGPT's search features, which surface websites[2]. OpenAI's documentation for its web search tool says the model chooses whether to search based on the prompt[3], and OpenAI says sites opted out of its OAI-SearchBot crawler will not be shown in ChatGPT search answers, though they can still appear as navigational links[2].
Does ChatGPT always search the web?
Two different systems sit behind one answer box. The first is what the model learned in training: text it took in up to a fixed date, its knowledge cut-off. If your number changed last year, or you dropped a service, an answer written from training alone may still be working from the old version of you.
OpenAI publishes a knowledge cut-off date for each of its three flagship models. On 25 September 2026 those three sat between 20 April and 18 May 2026, so an answer written from training alone can be months behind. [1]
The second is live search. OpenAI's documentation for its web search tool says the model decides for itself whether to search, based on the prompt, and that the query goes to the search tool, which returns a response based on top results[3]. With a reasoning model it goes further: the model can run searches as part of its chain of thought, analyse the results and decide whether to keep searching[3].
OpenAI's web search documentation says a response carries inline citations for URLs found in the results, and a separate sources list of every URL the model consulted, which often runs longer than the citations shown. [3]
So in the tool OpenAI documents, the pages behind an answer often outnumber the links under it. A page can inform what gets written about your trade without ever showing up as a link a reader can click.
Which OpenAI bot does what?
OpenAI's crawler page lists four user agents[2]. One of them, OAI-AdsBot, only visits pages submitted as ads on ChatGPT[2], so this guide covers the other three, which do three different jobs. Your robots.txt file names each one separately, so blocking one is not blocking the others.
| Crawler | What OpenAI says it is for | What blocking it does |
|---|---|---|
| OAI-SearchBot | Surfacing websites in search results in ChatGPT's search features[2] | An opted-out site is not shown in ChatGPT search answers, though it can still appear as a navigational link[2] |
| GPTBot | Crawling content that may be used in training OpenAI's generative AI foundation models[2] | Indicates the site's content should not be used to train those models[2] |
| ChatGPT-User | Certain user actions in ChatGPT and Custom GPTs, when someone asks a question[2] | Nothing in search: OpenAI says it is not used to determine whether content may appear in search[2] |
The distinction that matters to a business is that training and answering are separate decisions. Blocking GPTBot keeps your words out of future models. On OpenAI's own account it does not change whether ChatGPT search can fetch and cite you today. Blocking OAI-SearchBot does[2].
Two details are easy to miss. For search results, OpenAI says it can take about a day from a site's robots.txt update for its systems to adjust[2], and that because ChatGPT-User visits are started by a person, robots.txt rules may not apply to it[2]. OpenAI publishes the current IP ranges for all three[2]. The wider crawler table, covering Anthropic, Perplexity and Google as well, is in the GEO guide.
How does it handle ‘near me’?
OpenAI does not publish a method for choosing local businesses. What it does document sits on the developer side. In its web search tool, an API request can carry an approximate user location, a two-letter country code plus city, region and time zone, to refine search results based on geography[3]. That page documents the API tool, not the ChatGPT app.
None of that is a signal you send. What sits on your side of the line is whether your pages say, in text a crawler can read, which suburbs you actually serve. A service area that exists only inside a map graphic is not text.
OpenAI's crawler and web search documentation sets out how a search happens and how citations come back. It does not set out how one source is chosen over another[2][3]. Anyone reciting the ranking factors for ChatGPT is describing something OpenAI has not published.
What about shopping results?
ChatGPT can show products, and OpenAI documents how a merchant's catalogue gets in. It is a structured product feed shared with OpenAI, carrying titles, descriptions, images, price and availability, uploaded as a file and topped up through an API during the day[4]. The feed specification asks for a real brand and seller name rather than placeholders[5].
OpenAI's commerce documentation says onboarding product feeds in ChatGPT is currently available to approved partners, and points merchants at an application form. [4]
For a plumber, a salon or a tattoo studio, that settles one question: the documented feed is for products, and onboarding is open to approved partners only[4]. The way in is the ordinary one, a page OAI-SearchBot is allowed to fetch and can read as text[2].
What can I actually influence?
Six checks. Five of them are a browser tab or an edit you make yourself, and one, your robots.txt, is a decision about who you let in.
Watch what ChatGPT lists as Sources
In a new chat, ask for your trade and your suburb, and leave your business name out of it. Two things are worth reading: which businesses it names, and which pages it lists under Sources. Run the same question again a few days later, because each answer is written fresh rather than read off a standing list.
Check OAI-SearchBot is allowed
Open your own address followed by /robots.txt. If OAI-SearchBot is disallowed, OpenAI says your site will not be shown in ChatGPT search answers, except as a navigational link[2]. Allow it, then give it about a day to register[2].
Find your suburbs in the page source
Open the page and ask your browser for the source. Search that text for your phone number, the suburbs you serve and a service name. Words on screen but missing from the source are built by JavaScript or sitting in a graphic, and the GEO guide covers which crawlers JavaScript shuts out.
Put the answer in text
Hours in a photo of the shop door, a service area drawn on a map image, a price list inside a graphic: none of that is text a crawler reads. Write it as words on the page, under a heading that matches the question.
Make the facts agree
Name, phone, suburbs and hours, matching on your site, your Google Business Profile, your socials and every directory that lists you. Where OpenAI documents a search, the model consults a list of URLs rather than one page[3], and some of those pages will not be yours.
Ask ChatGPT to describe you
Give it your business name and read back what it says about your hours, your services and the area you cover. Where a detail is out of date, search the web for it and find the page still publishing that version. It may be an old directory listing, a social profile or your own site.
The free audit runs thirty weighted checks live in your browser against our own site, then scans your own site at our end, including crawler access, structured data and whether your pages are readable without JavaScript.
What did our own site teach us?
Our own robots.txt names OAI-SearchBot, ChatGPT-User and GPTBot, and allows all three. Its comment says why they are separate calls: whether a model may train on our words is a different decision from whether an assistant may fetch and cite us in an answer. Yours can land differently. The point is that it should be a decision rather than a default nobody looked at.
We have paid for the readability point ourselves: before its August 2026 rebuild, pixelategroup.com built its navigation and footer with JavaScript, and the GEO guide sets out what a crawler received instead.
Our own name has taught us something too: Google Search Console has shown pixelategroup.com in results for a different company's brand sitting one letter from our own, and the AEO guide carries the figures and the lesson.
Can anyone get a business named?
No, and it is worth being blunt. OpenAI publishes which crawler does what, when a search is triggered and how citations come back. It does not publish how one business is picked over another[2][3], and nobody outside OpenAI can see that. A service promising you a mention in ChatGPT's answers is selling control it does not hold.
No one controls which business an assistant names. The reasons yours gets ruled out are the part you can fix: a crawler locked out at the door, an answer that exists only as a picture, hours that disagree with your Google Business Profile, a page no search brings back. Our SEO, AEO and GEO page sets out that work, and the AEO guide covers how one passage gets quoted.
Questions and answers.
Should I block GPTBot?
That is a decision about training, not about being named. OpenAI says disallowing GPTBot indicates your content should not be used to train its foundation models, while OAI-SearchBot is the crawler that surfaces sites in ChatGPT's search features[2]. You can allow one and block the other. Ours allows both, and the robots.txt file says why.
How fast does a robots.txt change take effect?
About a day, on OpenAI's own account, before its search systems pick up a change to your robots.txt[2]. That is the permission, not the answer. What ChatGPT says about you also depends on when your pages were last fetched, and, where it answers from training instead of a search, on that model's knowledge cut-off date[1].
Why does ChatGPT have my old phone number?
Most likely it answered from the model rather than from a live search, so it is working from text gathered before that model's knowledge cut-off, a date months behind today for each of OpenAI's flagship models[1]. Anything that changed since reaches it only through a live search, and a live search can still find an old page somewhere carrying the old number.
Can I submit my business to ChatGPT?
OpenAI documents a product feed for commerce, with onboarding currently available to approved partners who apply[4], and advertising on ChatGPT, whose landing pages a separate bot, OAI-AdsBot, may visit to check them against OpenAI's policies[2]. For a service business that wants to be named in an answer, it documents no submission route. Everything else runs through the ordinary web, on a page OAI-SearchBot may fetch and can read as words[2]. The free audit checks both of those on your site.
Is this the same thing as GEO?
This is GEO applied to one assistant. GEO is the wider job of being visible to ChatGPT, Claude, Perplexity, Gemini, Copilot and Google's AI Overviews, and the GEO guide covers all six, with a crawler table for OpenAI, Anthropic, Perplexity and Google, plus llms.txt and the JavaScript problem. This guide stays inside ChatGPT, because OpenAI documents its own mechanics in unusual detail.
Sources.
- Models, with knowledge cut-off dates, OpenAI. Read 28 September 2026.
- Overview of OpenAI Crawlers, OpenAI. Read 28 September 2026.
- Web search tool (OpenAI API), OpenAI. Read 28 September 2026.
- Agentic commerce: get started with product feeds, OpenAI. Read 28 September 2026.
- Product feed specification, OpenAI. Read 28 September 2026.
See what ChatGPT can read on your site.
Before you ask ChatGPT about your trade and suburb, watch thirty weighted checks run in your browser against our own site, so you can see what each one tests. We then scan your own site at our end and email you the report. There is no card, and no call unless you want one.