Marketing

Indexed but Invisible: Why ChatGPT’s Labrador Index Finds Your Page but Doesn’t Cite It

Key Takeaways
  • Crawled, indexed, retrieved, listed, cited and recommended are different outcomes.
  • Labrador is a researcher-observed codename, not a public OpenAI product with documented optimization rules.
  • OAI-SearchBot—not GPTBot—is the crawler OpenAI identifies for ChatGPT search visibility.
  • Independent research indicates that ChatGPT often evaluates a page through a title and short stored snippet before it ever opens the full page.
  • A clear opening helps retrieval, but original evidence, relevance and source quality make a page more useful to cite.
  • ChatGPT can use different retrieval routes depending on the mode, account and question, so one prompt run is not a stable ranking report.
  • No special file, schema type or wording pattern can guarantee inclusion or citation.

Being available to ChatGPT does not guarantee that your page will be cited. A page can be crawlable, stored in a search index, retrieved for a prompt and listed in the Sources panel without earning a visible citation in the answer.

That distinction is the practical lesson behind recent research into an internal search index labeled Labrador. The name was observed by independent researchers in ChatGPT’s server-side result data; OpenAI has not published Labrador documentation, ranking factors or a submission tool. What OpenAI does officially confirm is that OAI-SearchBot is used for ChatGPT search, while GPTBot is associated with possible model training.

The short answer: To improve ChatGPT citation visibility, make your pages accessible to OAI-SearchBot, represent each page clearly in its title, H1 and opening text, answer commercially relevant questions, publish evidence worth citing and measure repeated prompts over time. None of these actions guarantees a citation, but each removes a different point of failure.

What is ChatGPT’s Labrador index?

Labrador is the name independent researchers found in a result_source field attached to some ChatGPT search results. The field reportedly appeared between May and July 2026 alongside other source labels, before disappearing from the exposed data on July 21.

Peec’s September 2026 investigation describes Labrador as a family of OpenAI-controlled indexes covering the general web and several specialized content types. Separately, RESONEO’s retrieval study identified Labrador as an in-house index rather than a renamed Google or Bing feed.

These are significant findings, but they are still reverse-engineering results. OpenAI’s public crawler documentation does not use the Labrador name or explain how pages in any internal index are scored. A reliable article therefore needs to separate three levels of evidence:

Evidence levelWhat we can reasonably sayWhat we cannot claim
Confirmed by OpenAIOAI-SearchBot supports ChatGPT search; GPTBot supports potential model training; the controls are independent.OpenAI has not published Labrador ranking factors or a Labrador submission console.
Observed by independent researchersThe Labrador label appeared in result data, ChatGPT used multiple retrieval pipelines, and pages moved through several visibility stages before citation.A 2026 experiment does not prove that the same pipeline mix or behavior applies to every user today.
Still unknownPublishers can test crawl access, page representation and citation outcomes.The exact ranking model, weighting, refresh schedule and reason one candidate beats another remain undisclosed.

Why indexed does not mean cited

Search visibility is a funnel, not a switch. The URL must survive several decisions before a user sees it in a ChatGPT answer.

StageWhat it meansWhat can go wrong
CrawlableOpenAI’s search crawler can request the URL.Robots rules, CDN security or server errors block access.
Indexed or storedA representation of the page is available to a retrieval system.The page has not been discovered, has an old stored copy or is represented poorly.
RetrievedThe page is selected as a candidate for a particular prompt or fan-out query.The content does not match the user’s need, wording, audience or constraints.
ListedThe URL appears somewhere in the answer’s Sources panel.It remains buried under “More” and contributes little visible exposure.
CitedThe answer visibly attaches the page to a claim.Another source supports the claim more directly or credibly.
RecommendedThe answer actively presents the brand or product as a suitable option.A citation may verify a fact without endorsing the brand.

ChatGPT visibility funnel showing crawlable, indexed, retrieved, listed, cited and recommended stages.

RESONEO examined 1,249 real ChatGPT answers, approximately 88,000 search results and 26,900 distinct pages. In that dated dataset, far more URLs were retrieved than promoted or cited, and only a small fraction were opened as full pages. The study also found that most retrieved pages remained low in the Sources panel instead of receiving a citation in the answer text. These figures describe one measured corpus, not a universal ChatGPT conversion rate, but they demonstrate why “our page was found once” is not the same as sustained visibility.

Seven checks for pages that are indexed but invisible

1. Allow the correct OpenAI search crawler

Start with the only control OpenAI documents directly. OpenAI identifies OAI-SearchBot as the crawler used to surface websites in ChatGPT search. The same documentation says sites that opt out will not appear in ChatGPT search answers, although a navigational link may still appear.

A basic configuration for a site that wants search visibility is:

User-agent: OAI-SearchBot
Allow: /

A publisher can allow search while declining potential model-training use because OpenAI treats the controls independently:

User-agent: OAI-SearchBot
Allow: /

User-agent: GPTBot
Disallow: /

Do not stop at the text file. Confirm that the page returns a successful response to legitimate OAI-SearchBot requests and that your CDN, firewall or bot-protection layer is not rejecting OpenAI’s published IP ranges. OpenAI says robots.txt changes may take about 24 hours to be reflected in its systems, but it does not promise when an individual URL will be recrawled, indexed or cited.

2. Check how the page represents itself before the full page is read

One of the most useful findings from the Labrador research is that retrieval often works from a compact representation rather than a fresh read of the entire page. In its July 2026 dataset, RESONEO found that Labrador results commonly contained the page title and a snippet capped just beyond 200 characters. The snippet was generally anchored around the H1 and the beginning of the rendered body, rather than the meta description.

This does not create a universal “200-character ranking factor.” It does create a sensible editorial test:

  • Does the page have one descriptive H1?
  • Does the first visible paragraph answer the page’s main question?
  • Do category labels, dates, banners or image text crowd out the answer?
  • Would the title and opening still make sense if the rest of the page were temporarily unavailable?

A weak opening delays the answer:

AI search is changing rapidly, and businesses everywhere are looking for new ways to stay ahead.

A useful opening identifies the problem immediately:

A page can appear in ChatGPT’s retrieved sources without earning an inline citation because retrieval and citation are separate selection stages.

Continue writing a strong meta description for Google and other retrieval routes. The research suggests Labrador may not use it for the stored snippet, while scraped search results can. Optimizing the opening does not replace conventional SEO; it protects the page across more than one retrieval path.

3. Match the question a buyer actually asks

A page is retrieved for a prompt, not for the abstract idea of “visibility.” If your page explains a broad category while the user asks for a tool with a specific integration, budget or location, a more precise page may be selected.

Compare these prompts:

  • What is AI content software?
  • Which AI content platforms publish directly to WordPress?
  • What is a practical AI content workflow for a five-person marketing team?
  • Which tools combine content creation, publishing and performance tracking?

They belong to the same topic but represent different decisions. Build pages around real sales questions, support queries, comparison requests and use cases. Do not manufacture dozens of near-identical pages by swapping one phrase; create a distinct page only when the user’s task and the answer genuinely change.

4. Give the answer a reason to cite you

Clear formatting can help a system understand a page. It does not make a derivative page the best source.

A citation-worthy page usually offers at least one asset that is difficult to replace:

  • Original research with a transparent sample and methodology
  • A first-hand test with dates, conditions and limitations
  • An official product specification, policy or dataset
  • A worked example that shows the calculation or decision process
  • A maintained comparison based on explicit criteria
  • An expert explanation that clarifies a commonly misunderstood issue

RESONEO observed that pages opened for a full read were much more likely to receive a citation than pages merely present in the retrieval set. It also found that opened pages often included official, regulatory or primary material. That is correlation within the study, not proof of a ranking factor, but the editorial lesson is sound: publish the evidence another writer would need to support a claim.

This also supports Google performance. Google’s people-first content guidance asks whether a page provides original information, substantial analysis, clear sourcing and demonstrable expertise instead of simply rewriting existing results.

5. Make the evidence easy to extract and verify

Useful content should remain understandable after the page is reduced to text. Use descriptive headings, short answer-first paragraphs, properly labelled tables, ordinary lists and links placed beside the claims they support.

For important claims, include:

  • The name of the original source
  • The date or period the evidence covers
  • The population, sample or scope
  • The relevant limitation
  • A direct link to the supporting page

Keep structured data accurate and consistent with the visible page, but do not present schema as a Labrador shortcut. RESONEO observed that JSON-LD was removed from the version passed during an on-demand page read, while noting that its experiment could not determine whether the index uses structured data earlier in the process. For Google’s AI features, Google separately states that no special AI file or schema type is required. Neither finding supports claims that adding FAQ schema automatically earns a ChatGPT citation.

6. Treat freshness as a content property, not a date-change tactic

An updated page and an updated copy inside a retrieval system are not necessarily the same thing. RESONEO’s tests distinguished a short index representation from a separate cache used for fuller page reads, and found that stored copies could persist after the origin page changed.

Publishers cannot force a Labrador refresh, but they can reduce ambiguity:

  • Keep the canonical URL stable when the underlying topic remains the same.
  • Update statistics, screenshots and product details when they materially change.
  • Show an accurate “last updated” date only after a substantive review.
  • Link important updates from relevant pages so crawlers can rediscover them.
  • Correct outdated claims rather than layering a new paragraph over them.

Do not change a date merely to look fresh. Google explicitly warns against changing dates when content has not substantially changed, and users lose trust when the visible timestamp does not match the evidence.

7. Measure citation visibility, not a single screenshot

ChatGPT does not provide a fixed public ranking position for a prompt. The answer can change with the wording, search mode, account, location, time and retrieval route.

RESONEO’s August follow-up found materially different source mixes between free Think and paid thinking modes for the same test set. Because the internal source label was removed in July, even researchers now have less direct visibility into which pipeline supplied a result. Treat every run as an observation, not a permanent rank.

For every monitored prompt, record:

  • Exact prompt wording and version
  • Date, country, language and ChatGPT mode
  • Whether the brand was absent, mentioned or recommended
  • Whether the owned site received an inline citation
  • Which third-party pages and competitors were cited
  • Whether the description of the brand was accurate

Repeat the same commercially relevant prompts on a consistent schedule. Look for patterns across a group of prompts and runs rather than reacting to one answer.

A 30-minute ChatGPT visibility audit

  1. Check crawler access: inspect robots.txt, the page response and any CDN or firewall rules affecting OAI-SearchBot.
  2. Read the search representation: copy the title, H1 and opening paragraph into a blank document. Confirm they explain the page without surrounding context.
  3. Define the prompt: write the real question a prospect would ask before needing this page.
  4. Find the unique evidence: identify the statistic, test, example, policy or expert explanation that makes your page more than a summary.
  5. Run repeated checks: test the same prompt more than once and record the mode and date.
  6. Inspect the sources: distinguish a URL buried in the Sources panel from an inline citation and from a recommendation.
  7. Choose the correct fix: solve the stage that failed instead of rewriting the entire page.
What you observeLikely problem areaBest next action
The URL never appears in sourcesCrawl access, discovery, prompt relevance or source coverageVerify OAI-SearchBot access, internal discovery and alignment with the prompt.
The URL appears under “More” but is not citedThe page was retrieved but another source supported the answer betterStrengthen the direct answer, evidence, authority and claim-level sourcing.
The brand is mentioned but a third party is citedChatGPT recognizes the entity but prefers external validationPublish primary product facts and earn accurate coverage on relevant third-party sites.
The answer cites old informationA stale index representation, cached page or stale third-party sourceUpdate the canonical page, correct prominent external profiles and monitor subsequent runs.
The brand appears only in some modesDifferent retrieval routes or search budgetsTrack each relevant mode separately instead of combining everything into one score.
The site is cited but the brand is not recommendedThe page supports a fact but does not establish product fitCreate evidence for the target use case, audience and buying constraint.

Diagnostic guide for fixing missing ChatGPT sources, citations and recommendations.

What not to do for Labrador visibility

  • Do not treat GPTBot as the search crawler. OpenAI assigns that role to OAI-SearchBot.
  • Do not assume a Bing ranking guarantees ChatGPT visibility. Independent research indicates that ChatGPT can blend its own index with several external routes.
  • Do not call llms.txt an OpenAI requirement. You may use it for other workflows, but OpenAI does not list it as a condition for ChatGPT search inclusion.
  • Do not stuff the first 200 characters. Write a useful answer for a reader, not a compressed block of repeated keywords.
  • Do not hide text for crawlers. Hidden claims create trust and accessibility risks and can expose the page to manipulation or spam concerns.
  • Do not claim schema guarantees citations. Valid structured data is good site hygiene, not proof of Labrador preference.
  • Do not manufacture research or freshness. A named methodology and honest limitation are more valuable than an unsupported “we tested” claim.
  • Do not report retrieval as citation. A page in a source list has not necessarily been shown prominently or attached to a claim.

Can small websites compete in the Labrador index?

Independent research suggests Labrador is not limited to publishers with formal OpenAI licensing agreements. RESONEO reported that partner and non-partner sites could receive the same basic result format within the observed index. That does not mean every site has equal visibility; it means a licensing deal was not shown to be a prerequisite for inclusion in the measured corpus.

A smaller site is most competitive when it becomes the best available source for a narrow, commercially meaningful question. A page based on a real test, an original dataset or deep operational experience can be more useful than a generic article on a larger domain.

The durable strategy is therefore not “optimize for Labrador” in isolation. Build pages that are technically accessible, useful in Google, clear in short-form retrieval and strong enough to remain credible when a model opens the full page.

How Minineo can support the audit

No external tool can show an official Labrador rank because OpenAI does not publish one. Minineo can still help with the parts a brand can measure: checking technical and structural barriers, monitoring a stable set of prompts, separating mentions from citations and comparing how competitors appear over time.

Start with Minineo’s free AI Visibility Score Calculator to identify page-level issues worth reviewing. Then validate the business outcome with repeated prompts and actual citation tracking. A readiness score is a diagnostic, not proof that the page is inside Labrador or guaranteed to be cited.

Frequently asked questions

Is Labrador an official OpenAI product?

No public OpenAI documentation currently presents Labrador as a product or names its ranking rules. The term comes from independent researchers who observed it in ChatGPT result data. Treat Labrador as a well-supported research finding, not an official optimization specification.

How do I submit my website to Labrador?

OpenAI does not provide a public Labrador submission form. The documented action for ChatGPT search eligibility is to allow OAI-SearchBot and avoid blocking its published IP ranges. Normal discoverability practices such as internal linking, stable URLs and accessible text remain sensible, but indexing and citation are not guaranteed.

Does allowing GPTBot improve ChatGPT search visibility?

OpenAI documents GPTBot for potential foundation-model training and OAI-SearchBot for search. You can allow OAI-SearchBot while blocking GPTBot. Do not use GPTBot access as evidence that a page is eligible for ChatGPT search.

Does Google ranking still matter for ChatGPT?

It can. The independent studies found that some ChatGPT modes relied heavily on Google-derived results while others leaned more on OpenAI’s own index. The source mix changed by mode and over time, so strong conventional SEO and direct OpenAI crawl access should be treated as complementary rather than competing strategies.

Does the meta description matter to Labrador?

RESONEO’s measured Labrador snippets were usually drawn from the rendered page body rather than the meta description. Other retrieval routes may use the meta description, so keep it accurate and persuasive while also improving the visible H1 and opening paragraph.

Does schema markup increase ChatGPT citations?

There is no official evidence that a particular schema type guarantees ChatGPT citations. Use valid structured data when it accurately represents the visible content and supports normal search features, but do not substitute markup for a clear answer, original evidence or credible sourcing.

How long does it take to appear in ChatGPT search?

OpenAI says robots.txt changes may take roughly 24 hours to affect its systems. It does not publish a guaranteed recrawl, indexing or citation timeframe for individual pages. Measure over several weeks and distinguish technical access from citation performance.

Is being listed in Sources the same as being cited?

No. A URL can appear in the Sources panel without receiving an inline citation or visible prominence. Track retrieved or listed pages separately from citations, recommendations and referral traffic.

Fix the visibility funnel, not the codename

Labrador changes how we understand ChatGPT search, but it does not create a secret checklist that replaces SEO. Its real value is diagnostic: it shows why crawl access alone is insufficient and why a page may be technically present yet practically invisible.

Start with the stage you can verify. Allow the correct crawler. Make the page’s purpose unmistakable. Publish evidence another answer would need. Keep important facts current. Then track whether the brand progresses from absent to mentioned, from mentioned to cited and from cited to recommended.

That is slower than chasing a supposed ranking trick. It is also far more likely to produce visibility that survives the next pipeline change.

Sanjay Negi

Sanjay Negi writes about AI-powered SEO, content automation, and practical strategies to improve organic visibility across Google and AI search. He shares hands-on insights from working with WordPress, AI tools, and scalable content workflows, with a focus on real-world results and platforms like Minineo.

More posts by Sanjay Negi →