ZenWeb - Blog - Crawl Budget: Why Google Ignores Some of Your Pages

Crawl Budget: Why Google Ignores Some of Your Pages

Jian Tat Lee
August 25, 2026

Share this post:

Crawl Budget: Why Google Ignores Some of Your Pages
TL;DR: Crawl budget is the set of URLs Googlebot will crawl on your site in a given period. Two things decide it: how fast your server responds (crawl capacity) and how much Google wants your pages (crawl demand). Most small Malaysian sites never hit the limit. Large or messy sites do — and Google quietly skips their less important pages.

You publish a page, wait a week, and it still isn’t in Google. No penalty, no error — it simply never got crawled. For big or cluttered sites, this happens every day, and that’s usually why.

Crawl budget is not a ranking factor you tune for a boost. It’s the gatekeeper that decides which of your pages Google even looks at. Get it wrong on a Malaysian e-commerce store, property portal, or classifieds site with thousands of filter URLs, and Google burns its time on junk while your money pages wait. This guide covers what it is, whether it matters for your site, where it leaks, and how to fix it. The video below, from Google’s own team, sets the scene in a few minutes.

Crawl Budget & the Crawl Stats Report, Explained by Google

Source video: Crawl Budget and the Crawl Stats report — Google Search Console Training

1. What Crawl Budget Actually Means

Quick Answer: Crawl budget is the set of URLs Googlebot can and wants to crawl on your site in a given period. Google sets it from two levers — crawl capacity (how many requests your server handles without slowing down) and crawl demand (how important and fresh Google thinks your pages are).

Before a page can rank, it has to be crawled, then indexed. Crawling is Googlebot fetching the page; indexing is Google deciding to store it. Crawl budget governs that first step. If Googlebot never spends the budget to fetch a URL, nothing after it can happen — the page is invisible. Our guide on how Google crawls and indexes your website walks through the full pipeline.

One detail trips people up: it’s set per hostname, not per company. So shop.example.my and www.example.my get separate budgets. Google devotes only so much time to any one host, so a site that wastes that time on low-value URLs leaves its important pages under-crawled. It is not a ranking signal itself — but it’s the plumbing behind almost every other part of how SEO actually works.

Key takeaway: Crawl budget = capacity (server speed) × demand (page importance), measured per hostname. It’s not a ranking factor, but no page ranks until Googlebot spends the budget to crawl it.

2. Does Crawl Budget Even Matter for Your Site?

Quick Answer: For most Malaysian small-business sites, no. If your pages get crawled the same week you publish them, it’s not your problem. It starts to matter once a site runs into the thousands of URLs, changes often, or shows a growing pile of “Discovered – currently not indexed” pages in Search Console.

Google is blunt about this. Its own crawl budget documentation says most sites don’t need to worry. If your pages are usually crawled the day you publish them, keeping your sitemap current and checking index coverage is enough. The advanced guidance is aimed at large sites (1 million+ URLs) and medium sites (10,000+ URLs) that change daily.

So the honest answer depends on your site’s size and mess, not your ambition. Here’s a rough guide to where you sit.

Does Crawl Budget Matter? A Site-Size Guide
Guide to whether crawl budget is a concern based on site size in URLs, the usual symptom, and the recommended action.
Site size (URLs)Does it matter?Usual symptomWhat to do
Under 500 (most SMEs)RarelyPages usually crawled within daysKeep the sitemap updated; nothing else
500–10,000SometimesSome new pages slow to appearTidy internal links and prune thin pages
10,000+ & changes dailyYes“Discovered – not indexed” keeps growingActively manage it
1,000,000+AlwaysDeep pages never crawled at allRun a full crawl-budget programme

Illustrative bands, based on Google Search Central’s published thresholds. Numbers are rough classifications, not exact cut-offs.

If Search Console shows pages stuck as “Discovered – currently not indexed” or crawled but not indexed, that’s your signal that crawling or page quality needs attention.

Key takeaway: Don’t chase this on a 40-page brochure site. Chase it when you’re past a few thousand URLs, publish often, or watch “Discovered – not indexed” climb month after month.

Not sure whether crawl budget is holding your site back?

ZenWeb reads your Crawl Stats and index coverage and tells you straight — see how our SEO service digs into technical health →


3. Where Your Crawl Budget Leaks Most

Quick Answer: Crawl budget rarely runs out because a site is too big. It runs out because Googlebot spends its time on URLs that shouldn’t exist — filter combinations, duplicates, redirect chains, and dead pages. On the Malaysian sites we audit, the single biggest drain is filter and sort parameters spawning near-infinite URLs.

When we run technical audits, the wasted crawls fall into a handful of repeat offenders. The chart below shows the rough split across Malaysian SME sites we’ve reviewed.

Where Crawl Budget Gets Wasted on Malaysian SME Sites
Approximate share of wasted Googlebot crawls by cause across audited Malaysian SME sites: filter and sort parameters, duplicate pages, redirect chains, soft 404s and thin pages, tracking-parameter URLs, and slow pages.
Cause of wasted crawlShare of wasted crawlsShare
Filter & sort URL parameters
34%
Duplicate & near-duplicate pages
22%
Redirect chains & loops
16%
Soft 404s & thin/expired pages
14%
Tracking & session-ID URLs
9%
Slow pages cutting crawl rate
5%

Aggregated from ZenWeb-managed technical audits across Malaysian SME sites, 2024–2026. Shares are approximate and vary by platform.

Notice the pattern: almost none of this is “too much good content.” It’s duplicate content and messy URL structures generating pages nobody needs. Fix the leaks and the budget looks after itself.

Key takeaway: Over half of wasted crawls come from filter parameters and duplicates. You don’t buy more of it — you stop wasting the budget you already have.

4. How to Optimise Your Crawl Budget, Step by Step

Quick Answer: You optimise it by cutting the URLs Google shouldn’t crawl and speeding up the ones it should. In order of impact: consolidate duplicates, block junk URLs in robots.txt, return clean 404s, kill redirect chains, keep the sitemap tight, strengthen internal links, and make pages load fast.

These steps mirror Google’s own best practices, worked into the order we apply them on client sites. Do them top to bottom — the early steps remove the most waste.

  1. Consolidate duplicate content. Point near-identical pages at one canonical URL so Googlebot crawls a single version, not five. Start with the worst offenders flagged in your duplicate content audit.
  2. Block junk URLs with robots.txt. Filter, sort, and internal-search URLs that you never want indexed belong in robots.txt. Check first that you’re not accidentally blocking Google from pages you do want.
  3. Return a clean 404 or 410 for dead pages. A real 404 tells Google to stop crawling a removed URL. Soft 404s — pages that look empty but return 200 — keep getting recrawled and waste budget.
  4. Fix redirect chains. Every hop in a chain is an extra crawl. Collapse chains so each old URL points straight to the final destination with a single 301 redirect.
  5. Keep your sitemap tight and current. List only canonical, indexable URLs and add a <lastmod> date. A clean XML sitemap tells Google exactly what deserves a crawl.
  6. Strengthen internal links and hierarchy. Pages buried deep in the site get crawled less. A flatter, well-linked website structure pushes crawl demand toward your key pages.
  7. Make pages fast. When your server responds quickly, Google’s crawl capacity limit rises and it fetches more per visit. Chasing down slow page speed literally buys you more crawling.
Key takeaway: Optimising it is mostly subtraction — remove the URLs Google shouldn’t touch, then make what’s left fast and easy to reach.

5. Crawl-Waste Patterns by Platform

Quick Answer: Each platform wastes it in its own way. WordPress spawns thin tag and archive pages, WooCommerce explodes filter combinations, custom builds leak session IDs, and Shopify duplicates product URLs across collections. Knowing your platform’s default trap tells you where to look first.

The leaks in Section 3 don’t appear at random — they follow the CMS. Here’s where each common platform tends to burn budget and the quickest fix.

Common Crawl Traps by Platform
The most common crawl trap for each website platform and the quickest fix, covering WordPress, WooCommerce, custom or dynamic builds, and Shopify.
PlatformMost common crawl trapQuickest fix
WordPressThin tag, author, date and attachment pagesNoindex the archives you don’t use; disable attachment pages
WooCommerceLayered-nav filter combinations (colour + size + price)Block filter parameters in robots.txt; canonicalise
Custom / dynamicSession IDs and tracking params creating endless URLsStrip params server-side; set a clean canonical
ShopifySame product on multiple /collections/ pathsCanonicalise products to their /products/ URL

Patterns observed across ZenWeb-managed Malaysian client sites, 2024–2026.

If you run a store, the filter problem is usually the big one. A messy URL structure with stacked parameters can turn 500 products into hundreds of thousands of crawlable URLs.

Key takeaway: Find your platform’s default trap first — it usually accounts for most of the waste before you look at anything custom.

6. What Cleaning Up Crawl Budget Actually Changes

Quick Answer: When you cut wasted URLs, Google redirects the same crawl effort onto pages that matter. In practice that means more of your important pages crawled each day, fewer stuck in “Discovered – not indexed,” and new pages appearing in Google in days rather than weeks.

The table below tracks a representative cleanup on a mid-size Malaysian store — the shape we see repeat across similar projects.

Crawl Metrics Before and After a Cleanup
Crawl metrics before a cleanup and at week 4 and week 8 after, covering important pages crawled per day, share of crawl on important pages, pages stuck in discovered-not-indexed, and average days for a new page to be indexed.
MetricBeforeWeek 4Week 8
Important pages crawled / day~380~610~840
Share of crawl on important pages~45%~68%~82%
Pages in “Discovered – not indexed”~9,200~5,100~2,400
Avg days for a new page to index~18~7~3

Illustrative recovery pattern, aggregated from ZenWeb-managed Malaysian site cleanups, 2024–2026. Figures vary by site size and starting condition.

None of this changes your rankings overnight. But faster, cleaner crawling means your best work gets seen sooner — and every part of how Google ranks pages depends on Google seeing them first. It’s foundational to ranking a Malaysian business on Google.

Key takeaway: Cleaning it up doesn’t add rankings directly — it removes the delay between publishing and getting seen, which is often the real bottleneck.

Got thousands of pages stuck in “Discovered – not indexed”?

That’s a crawl budget cleanup, and we do them often — get a technical SEO audit from ZenWeb →


7. Crawl Budget Mistakes That Waste Googlebot’s Time

Quick Answer: The most common mistakes come from using the wrong tool for the job. People reach for noindex when they mean robots.txt, block pages they actually want ranked, or set canonicals that point Google at the wrong URL. Each one quietly costs you crawls.

These are the errors we see most often on Malaysian sites, and why they backfire:

  • Using noindex to save budget. Google still has to crawl a page to see the noindex tag, so it doesn’t save budget at all. If you never want a URL crawled, block it in robots.txt instead. See how noindex tags can block the wrong pages.
  • Blocking pages you want indexed. An over-broad robots.txt rule can hide money pages from Google entirely. Always test before you deploy — a stray line can block Google from your whole site.
  • Broken canonicals. A canonical pointing at the wrong URL, or every page canonicalising to the homepage, scatters crawl signals. Watch for these canonical tag mistakes.
  • Leaving soft 404s live. Empty pages that return a 200 status keep getting recrawled forever. Give removed pages a real 404 or 410.
  • Infinite spaces. Calendars, endless pagination, and “load more” URLs create traps Googlebot can crawl forever without finding anything new.
Key takeaway: robots.txt stops a crawl; noindex stops an index. Mixing them up is the number-one mistake here — noindex still spends the crawl you were trying to save.

8. How to Know Your Crawl Budget Is Working

Quick Answer: Check the Crawl Stats report in Google Search Console (Settings → Crawling). Healthy crawling shows steady or rising total crawl requests, most crawls returning a 200 status, fast average response times, and new pages moving out of “Discovered – not indexed” and into the index.

You don’t need fancy tools to monitor this. Google Search Console gives you the whole picture for free:

  • Crawl Stats report. Watch total crawl requests, average response time, and the breakdown by response code. A wall of 404s or 5xx errors means budget is leaking — check the detail against your Search Console crawl errors.
  • Page indexing report. Track the count of pages not getting indexed. If “Discovered – not indexed” is shrinking, your cleanup is working.
  • Watch for slippage. A sudden jump in uncrawled URLs, or pages dropping out of the index, is your early warning to audit again.
Key takeaway: The Crawl Stats and Page Indexing reports in Search Console are your dashboard. Rising crawl on 200-status pages and a shrinking “Discovered – not indexed” count mean the budget is landing where it should.

9. Conclusion: Make Every Crawl Count

Crawl budget isn’t a dial you turn up. It’s the outcome of a tidy, fast site that points Google at the pages that matter and keeps it away from the ones that don’t. For most small Malaysian businesses, a current sitemap is enough. For stores, portals, and any site running into the thousands of URLs, it becomes the difference between publishing and being seen.

At ZenWeb, it’s part of the technical SEO groundwork we lay before chasing rankings — because the best content earns nothing while it sits uncrawled. Fix the leaks first, and everything downstream works harder.

Pages published but not showing up in Google?

Book a free 30-minute strategy session. We’ll read your Crawl Stats, find where Googlebot is wasting its time, and map the fixes that get your important pages crawled and indexed faster.

Get your free SEO audit →


10. Frequently Asked Questions

1. What is crawl budget in simple terms?

Crawl budget is how many of your pages Google will fetch in a given period. It’s set by how fast your server responds and how much Google wants your content. If Googlebot spends that budget on junk URLs, your important pages wait longer to be crawled and indexed.

2. How do I check my crawl budget?

Open Google Search Console and go to Settings → Crawling → Crawl Stats. It shows total crawl requests, average response time, and status codes over 90 days. If Google isn’t crawling enough, our guide on getting a site crawled walks through the fixes.

3. Do subdomains and language versions have separate crawl budgets?

Yes. It’s set per hostname, so blog.example.my and www.example.my each get their own. This matters for international SEO setups — if you split languages across subdomains and use hreflang tags, each host is crawled on its own budget.

4. Does crawl budget affect mobile?

Google crawls with its mobile bot first, so your mobile site is what gets the budget. A slow or bloated mobile experience lowers crawl capacity, meaning fewer pages fetched per visit. Strong mobile SEO and fast mobile pages directly help crawling.

5. My page is crawled but not ranking — is that a crawl budget problem?

No. Crawling only gets a page seen; ranking is a separate contest. Once a page is crawled and indexed, it competes for SERP features like featured snippets, People Also Ask boxes, and voice search answers — which come down to content and relevance, not crawl budget.

Table of Contents

Table of Contents

See Also

Best Meta Ads for Equipment Rentals in Malaysia: Guide 2026

Best Meta Ads for Equipment Rentals in Malaysia: Guide 2026

Best Google Ads for Equipment Rentals in Malaysia Guide 2026

Best Google Ads for Equipment Rentals in Malaysia Guide 2026

Best SEO for Equipment Rentals in Malaysia: Guide 2026

Best SEO for Equipment Rentals in Malaysia: Guide 2026

Get A Free Proposal

Complete the form and our team will contact you to discuss your goals. Let’s grow your business.

Meowketing Specialist

Online

Today

Meow! 👋

We are Official Google Partner,
Ask us anything about Marketing!