ZenWeb - Blog - AI Crawlers: Should You Let GPTBot Read Your Website?

AI Crawlers: Should You Let GPTBot Read Your Website?

Jian Tat Lee
August 12, 2026

Share this post:

AI Crawlers: Should You Let GPTBot Read Your Website?
TL;DR: For most Malaysian businesses, yes. GPTBot only collects training data, so blocking it costs you nothing in ChatGPT visibility. The bots that decide whether ChatGPT, Gemini and Perplexity can quote you are different ones, and blanket “block all AI bots” rules switch those off by accident.

1. Introduction

The usual version of this question is a straight yes or no: block the AI bots, or let them in and hope for traffic. Both answers are wrong, and most sites have already answered it by accident anyway.

There is no single AI crawler. Four different jobs arrive under separate names, each with a different consequence when blocked, and a rule aimed at one routinely catches the others.

This guide covers which bots reach a Malaysian site, what blocking each one costs, what we find in real robots.txt files at audit, and the rule set we recommend. For the wider programme, see our SEO services or the ZenWeb home page. The walkthrough below is a fast primer.

Technical SEO for AI: Robots.txt, GPTBot & llms.txt Explained | 3.4. AEO Course by Ahrefs

Source video: Ahrefs on YouTube


2. Which AI Crawlers Actually Reach a Malaysian Site

Quick Answer: Four kinds of AI bot hit a typical Malaysian website: training crawlers that feed model weights, search crawlers that build the index behind AI answers, live fetchers triggered by a user’s question, and ad-safety checkers. Each announces itself with a different user-agent name.

The names matter because robots.txt matches on them. OpenAI alone documents four separate agents:

  • Training crawlers. GPTBot, ClaudeBot, CCBot and the Google-Extended token. They collect text that may train a model.
  • Search crawlers. OAI-SearchBot and PerplexityBot build the retrieval index AI answers are drawn from and cited against.
  • Live user fetchers. ChatGPT-User fires when someone asks a question that needs a page opened right now.
  • Ad checkers. OAI-AdsBot visits pages submitted as ChatGPT ad landing pages to confirm policy compliance.

Most of this is familiar ground if you know how Google crawls and indexes a website. What changed is the number of separate decisions you now have to make.

Key takeaway: “AI crawler” is not one thing. Four jobs arrive under different names, each needing its own decision.

3. What Blocking GPTBot Does and Does Not Do

Quick Answer: Blocking GPTBot tells OpenAI not to use your pages for model training. It does not remove you from ChatGPT search answers, because those come from a different crawler. Most owners block GPTBot believing they are opting out of ChatGPT entirely, and they are not.

OpenAI is explicit about the separation. Its documentation states that a webmaster can allow OAI-SearchBot in order to appear in search results while disallowing GPTBot. Two switches, two outcomes.

Google works the same way. Its crawler reference confirms that Google-Extended does not impact a site’s inclusion in Google Search nor is it used as a ranking signal.

So blocking training crawlers is a rights decision. Blocking search crawlers is a marketing decision, and a costly one if you make it by accident. The same split runs through what actually changes between AI SEO and traditional SEO.

Key takeaway: Blocking GPTBot costs nothing in ChatGPT citations. Blocking OAI-SearchBot removes you entirely. Confusing the two is the expensive mistake.

Not sure what your site is currently blocking?

We read your robots.txt, CDN rules and AI mentions together before recommending a single change. See what an AI visibility audit covers →


4. Crawler by Crawler: What Blocking Each One Costs

Quick Answer: Blocking a training crawler costs visibility nothing. Blocking a search crawler removes you from that engine’s answers. Blocking a live user fetcher means the one person who asked about you gets nothing. The table below maps each bot to its real consequence.

This is the reference we build every client rule set from, compiled from the operators’ own documentation rather than third-party lists.

What blocking each AI crawler costs
AI crawler tokens mapped to operator, job and the visibility consequence of blocking each one.
robots.txt tokenOperatorJobWhat blocking it costs you
GPTBotOpenAIModel trainingNo visibility loss
OAI-SearchBotOpenAIChatGPT search indexRemoved from ChatGPT search answers
ChatGPT-UserOpenAILive fetch on a user’s requestThe asking user gets nothing from your page
OAI-AdsBotOpenAIAd landing page checksOnly matters if you advertise in ChatGPT
Google-ExtendedGoogleGemini training and groundingNo Google Search loss; weaker Gemini grounding
ClaudeBotAnthropicModel trainingNo visibility loss
PerplexityBotPerplexityAnswer indexRemoved from Perplexity citations
CCBotCommon CrawlOpen dataset used by many modelsNo visibility loss; broad training opt-out

Source: compiled by ZenWeb from OpenAI, Google and Anthropic crawler documentation, July 2026.

Anthropic publishes its own guidance too, confirming that ClaudeBot honours standard robots.txt directives. Why a table like this earns citations is answered in formatting pages that LLMs can quote.

Key takeaway: Only three tokens in that table cost real visibility. The rest are training opt-outs with no marketing downside.

5. The Crawl-to-Referral Trade Nobody Priced In

Quick Answer: AI platforms crawl far more than they refer back. Cloudflare publishes the ratio per platform, and the gap between the heaviest crawler and a traditional search engine is enormous. That gap is the real argument for blocking training bots.

Search engines struck a bargain publishers understood: we read your page, we send you the reader. AI answers break that trade. Cloudflare now reports the gap publicly, noting that for the week of 19–26 June 2025 the ratios ranged from Anthropic’s 70,900:1 down to Mistral’s 0.1:1. It also cautions that native-app traffic often carries no referrer, so the ratios may overstate the gap.

A search crawler reads your page to send you a visitor. A training crawler reads it so nobody has to visit at all.

The figures move weekly and are worth checking on Cloudflare Radar’s AI Insights page before deciding. For a Malaysian SME the read is simple: refuse the bots that take without giving, keep the ones that can send an enquiry. We applied the same test to whether AI referral visitors actually convert.

Key takeaway: The crawl-to-referral gap justifies blocking training bots. It never justifies blocking bots that can still send you a customer.

6. What Malaysian SME robots.txt Files Actually Say

Quick Answer: Across the Malaysian SME sites we audit, most have no AI crawler rules at all, and a meaningful minority are blocking retrieval bots without knowing it. Almost none made a deliberate, documented decision either way.

We check this on every intake, alongside the basics in technical SEO made simple. The pattern has held for two years.

AI crawler rules found at audit intake
Share of audited Malaysian SME websites carrying each AI crawler configuration at engagement.
What we foundShare of sites
No AI crawler rules of any kind

68%

GPTBot blocked

19%

Retrieval bots blocked unintentionally

14%

Google-Extended blocked

9%

robots.txt missing or returning an error

7%

A written decision on AI access exists

4%

Source: ZenWeb client tracking, Malaysian SME technical audits, 2024–2026.

The 14% is the line that costs money. Nobody chose it; it arrived through a copied rule set, a plugin default or a CDN toggle. Start with checking whether robots.txt is blocking Google.

Key takeaway: Almost nobody decides this on purpose. Check what your file says before arguing about what it should say.

Worried you are in that 14%?

A crawler access check is the first item in our AI search work, before any content is rewritten. See the first 90 days of AI search fixes →


7. The CDN Setting That Blocks Bots You Never Named

Quick Answer: Your robots.txt is not the only gate. CDN and security platforms now ship AI bot controls that can be on by default, so a site can be blocking retrieval bots at the network edge while its robots.txt says nothing at all.

This catches experienced teams. The developer reads robots.txt, sees no AI rules, and calls the site open, while the CDN refuses requests before they reach the server. Four places to check:

  • The CDN or security platform. Find the AI bot or scraper control panel and read what the default actually does.
  • Your hosting or WordPress security plugin. Several ship bot blocklists that update without telling you.
  • Server-level rules. Old .htaccess or firewall entries copied from a forum thread years ago.
  • Your own server logs. The only honest test. OAI-SearchBot getting a 200 means you are open; a 403 means you are not.

Rendering is the silent fifth blocker. If your key content only appears after JavaScript runs, treat it as invisible, which is also the reasoning behind teaching Google and AI who your company is in plain server-rendered text.

Key takeaway: Robots.txt is a request; the CDN is a wall. Read your logs, because only the logs tell you which one is actually deciding.

8. AI Referral Traffic by Industry: Is It Worth Protecting?

Quick Answer: AI referrals are a small share of sessions for most Malaysian SMEs, but they carry unusually high intent. Considered-purchase sectors see the strongest enquiry rates, while high-volume retail sees the weakest.

Volume is the wrong measure. What matters is how many of those sessions become a real enquiry.

AI referral sessions by industry
Median monthly AI referral sessions and enquiries across tracked Malaysian client sites by industry.
IndustryOrganic sessionsAI referral sessionsShare of totalEnquiries / month
B2B manufacturing and supply2,6001435.5%5
Professional services4,2001764.2%6
Dental and medical clinics6,8002313.4%8
Property and renovation9,4002732.9%9
F&B and retail12,1002181.8%3

Source: ZenWeb client tracking, Malaysia, 2025–2026.

B2B suppliers get the smallest raw numbers and the highest share, because buyers researching a specification are the people asking an assistant instead of scrolling ten blue links. Reading these numbers starts with tracking AI citations across engines.

Key takeaway: AI referrals are small in volume and strong in intent. Judge them on enquiries per hundred sessions, never on session share.

9. The robots.txt Rule Set We Recommend

Quick Answer: For a lead-generating Malaysian SME site, allow every retrieval and user-fetch bot, then decide on training bots as a business question. Write the rules by name, never with a wildcard, and confirm the result in your server logs.

Setting AI crawler rules on a Malaysian SME site

Run these in order. The whole job is under an hour, and the last step is the one people skip.

  1. Read what you have now. Open your robots.txt, then your CDN’s bot settings. Write down what each one does before changing anything.
  2. Name the retrieval bots and allow them. Add explicit allow groups for OAI-SearchBot, ChatGPT-User and PerplexityBot so a later wildcard cannot catch them.
  3. Decide on training bots as a business call. If your content is your product, disallow GPTBot, ClaudeBot, CCBot and Google-Extended. Otherwise, leave them.
  4. Delete every wildcard AI rule. A blanket disallow aimed at unnamed AI agents produces the 14% in the audit data above.
  5. Confirm in your server logs. Wait a week, filter for each user agent, check the status codes. A 403 means something upstream is still refusing.
  6. Re-check after any platform change. Migrations, CDN switches and plugin updates all rewrite these rules silently.

Rules decide whether a bot can read the page; structure decides whether it gets cited, which is the argument in our 15-step AI SEO checklist. Whether a separate llms.txt file helps is a much smaller question.

Key takeaway: Name every bot explicitly and never use a wildcard. Then verify in the logs, because the file only tells you what you asked for.

10. What a Blanket Block Costs Over Twelve Months

Quick Answer: Blocking retrieval bots by accident is not a rounding error. Modelled at a 3% enquiry rate and a 20% close rate, a mid-sized Malaysian site gives up roughly twenty deals a year, most of them from high-intent research queries.

These figures are a projection, not a measurement, applying the referral shares above to three common site sizes.

Modelled annual cost of blocking retrieval bots
Modelled annual enquiries, deals and revenue forgone when retrieval bots are blocked, by site size.
Site sizeAI sessions lost / yearEnquiries lostDeals lostRevenue at RM 6,000 deal
Small (2,000 sessions / month)1,150357RM 42,000
Mid-sized (8,000 sessions / month)3,40010220RM 120,000
Established (20,000 sessions / month)7,90023747RM 282,000

Modelled projection using ZenWeb AI referral shares, 3% enquiry rate, 20% close rate.

Swap in your own deal value and the shape holds. At any realistic Malaysian ticket price, an hour of work returns more than most month-long content projects. On larger sites a single bad rule multiplies fast, which is the theme of scaling enterprise SEO past 1,000 pages.

Key takeaway: An accidental blanket block is a five-figure annual mistake, and it takes an hour to undo.

Want this checked properly, not guessed at?

Crawler access is a standard line item in every SEO engagement we run. Compare our SEO service tiers →


11. When Blocking Genuinely Makes Sense

Quick Answer: Block training crawlers when your content is the product you sell, when you licence it, or when it is proprietary research you do not want reproduced without attribution. Those are the honest cases, and none of them require blocking retrieval bots.

The cases where blocking is right:

  • Content is the product. Paid courses, subscription research, member libraries. Training on it undercuts what you sell.
  • You licence your data. If a publisher pays for your dataset, free training access weakens the deal.
  • Research you want cited by name. Blocking training while allowing retrieval keeps attribution and drops absorption.
  • Crawl load is hurting the server. A real problem on large sites, and closer to a crawl budget question than a policy one.

Most Malaysian SMEs fit none of these. If your website exists to make the phone ring, the content is marketing, and marketing wants to be repeated. Third-party mentions work the same way, as covered in whether directory citations feed AI answers and why AI keeps quoting forums instead of your site.

Key takeaway: Block training when your content is the product. If your content is the shop window, blocking is just closing the curtains.

12. How ZenWeb Handles This for Malaysian Clients

Quick Answer: We audit crawler access in week one of every SEO engagement, put the training decision in writing with the client, then re-verify in server logs each quarter and after any platform change.

ZenWeb is a Google Partner working with over 500 clients across Malaysia, Japan and Vietnam. Crawler access is a checklist item, never a separate invoice, and sits beside the reporting in benchmarking AI search share of voice:

  • A written access decision. Which bots are allowed, which are refused, and the reason for each.
  • Log verification, not file inspection. We confirm the status codes each bot actually receives.
  • Re-checks after every platform change. Migrations and plugin updates are where good rules disappear.

Ongoing work is scoped as we set out in what an AI SEO retainer should include. Entity work such as earning a Google knowledge panel or the honest case for Wikidata follows once access is settled.

Key takeaway: Treat crawler access as a documented decision that gets re-verified, not a file edited once and forgotten.

13. Conclusion

Quick Answer: Let the retrieval bots in, because they are the only ones that can send you a customer. Decide on the training bots separately, on commercial grounds. The costly error is not blocking GPTBot; it is blocking everything by accident.

Read your robots.txt, then your CDN settings, then your logs. Most Malaysian sites will find nothing configured at all, which is safer than it sounds and still worth setting deliberately.

Once access is clean, the work shifts to being worth quoting. That is a content and structure problem, and it is where the rest of our SEO and AI visibility work starts.


14. Frequently Asked Questions

1. Does blocking GPTBot remove me from ChatGPT?

No. GPTBot only collects training data. ChatGPT search answers are built from OAI-SearchBot’s index, and live page visits come from ChatGPT-User. Block GPTBot and allow the other two, and your pages can still be retrieved and cited in ChatGPT answers.

2. Will blocking Google-Extended hurt my Google rankings?

No. Google’s crawler documentation states plainly that Google-Extended has no effect on inclusion in Google Search and is not a ranking signal. It only governs whether your content trains Gemini models and grounds Gemini answers. Search and AI training are separate decisions.

3. Do AI crawlers actually obey robots.txt?

The major documented ones do, including GPTBot, ClaudeBot and PerplexityBot. Robots.txt is a request, not a lock, so unidentified scrapers may ignore it. Real enforcement happens at your server or CDN, which is why log checks matter more than the file itself.

4. Should a small Malaysian business block AI crawlers?

Usually not. For a business whose website exists to generate enquiries, AI referrals arrive with high buying intent and cost nothing to receive. The exception is when your content is the product you sell, such as paid research or a subscription library.

5. How do I check which AI bots are visiting my site?

Filter your server access logs by user agent for GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot and Google-Extended, then check the status codes returned. A 200 means the bot is being served; a 403 means something is refusing it before your content loads.

Ready to be the source AI engines quote?

Book a free 30-minute strategy session. We review your crawler access, your Google visibility and your competitors, then give you a 90-day plan with realistic enquiry targets.

Get my free strategy session →

Table of Contents

Table of Contents

See Also

E-E-A-T for AI: Prove Expertise Machines Can Verify

E-E-A-T for AI: Prove Expertise Machines Can Verify

Content Formats for AI Search: What Chatbots Prefer

Content Formats for AI Search: What Chatbots Prefer

Prompt Keyword Research: How Buyers Actually Ask AI

Prompt Keyword Research: How Buyers Actually Ask AI

Get A Free Proposal

Complete the form and our team will contact you to discuss your goals. Let’s grow your business.

Meowketing Specialist

Online

Today

Meow! 👋

We are Official Google Partner,
Ask us anything about Marketing!