Crawl Budget
I wasted months fixing server speed before realising that removing duplicate URLs is the fastest way to reclaim crawl budget.
Start here
- Check your Crawl Stats report in Google Search Console to see how many URLs Googlebot is hitting and how fast.
- Identify and remove duplicate content — wasted crawling on duplicates can make Googlebot spend less time on your site overall.
- Remove or fix soft 404s and long redirect chains; they eat budget with no indexing benefit.
- Keep XML sitemaps up to date with
<lastmod>tags to signal changed pages. - Monitor server response times only after cleaning up technical waste.
Plain-English take
Crawl budget is the number of pages Googlebot can and will crawl on your site within a given timeframe. It's a combination of two things: crawl rate (how fast Googlebot can fetch pages without overloading your server) and crawl demand (Google's interest in your content based on quality, freshness, and popularity). For a small blog with 50 pages, it's rarely an issue — Google can crawl all of them easily. But for an ecommerce site with 50,000 product pages, crawl budget becomes a bottleneck. Google won't crawl everything every day. If half of those pages are [duplicate](/duplicate-content/) or error pages, Google wastes slots and may not get to new content. The principle is simple: remove the dross first, then optimise the rest. I start every audit by checking the Crawl Stats report in Google Search Console. That tells me how many requests Googlebot made and how fast. Then I look at what pages it's actually crawling — often a mess of parameterised URLs and old categories. A [technical SEO](/technical-seo/) audit frequently reveals that the biggest win is not speeding up the server but cutting the crawl surface area.
When it actually matters
Crawl budget matters in three specific scenarios. First, large sites with 10,000+ pages. If you have 100,000 URLs but only 20,000 get indexed, budget is likely misallocated. Second, sites that update frequently — news sites, ecommerce with rolling stock, or blogs with daily posts. You want new content crawled within hours, not weeks. Third, sites with technical debt: long [redirect chains](/301-redirect/), [duplicate content](/canonical-tags/) from parameterised variants, or [endless pagination](/pagination/) without proper handling. Each extra page that shouldn't be crawled steals a slot from one that should. I've seen a site with 30,000 URLs but only 8,000 worth crawling — the rest were parameterised variants. After consolidating with canonical tags and removing junk, crawl efficiency doubled. But if your site is under 500 pages and you rarely update, crawl budget is a distraction. Focus on content quality and internal linking instead. Keep your [site structure](/website-structure/) clean and your sitemap lean.
What I got wrong
I used to think crawl budget was a direct ranking factor, like Google saying 'you have budget, you rank'. It's not. It's an operational constraint — if Google can't find your pages, they can't be indexed, which blocks ranking. But having more crawl budget doesn't boost rankings otherwise. The bigger mistake I made was prioritising server speed before cleaning up crawl waste. I spent weeks optimising TTFB and database queries, but Google was still wasting half its crawls on duplicate product pages and soft 404s. The fix was obvious in hindsight: remove the waste, then speed matters. I also left old, thin-content pages in my sitemap, thinking they didn't harm. They did — they signalled to Google that those pages were important, triggering unnecessary crawls. Now I audit sitemaps quarterly and only include pages I actually want indexed. Using [robots.txt](/robots-txt/) to block unimportant sections can also help, but it's not a silver bullet.
Next step
Quick answers
Does crawl budget matter for a small personal blog?
Usually no. If you have under 500 pages and update once a week, Google can crawl all your content easily. Focus on content quality and internal links instead. Crawl budget only becomes a real constraint on larger, faster-changing, or technically flawed sites.
How do I check if my site has a crawl budget problem?
Use the Crawl Stats report in Google Search Console. It shows total requests made by Googlebot, time spent, and average response time. Look for spikes on low-value URLs. Pair it with the URL Inspection tool to see when specific pages were last crawled.
Should I use robots.txt to block low-value pages?
Generally no. Disallowed URLs still consume budget because Googlebot tries to crawl them and gets blocked. Instead, use noindex tags for pages you don't want indexed, and remove or consolidate the rest. Blocking in robots.txt can prevent Google from seeing important content.
Sources
Primary documentation is linked directly. Anything commercial is marked nofollow.
- Google Search Central: Crawl budget management — Primary source for Google's definition and causes of wasted crawl budget.
- Google Search Central: Manage crawling and indexing of site URLs — Supporting documentation on controlling crawl and index behaviour through technical SEO signals.
- Google Search Central: Crawl stats report — Authoritative reference for measuring crawl activity in Google Search Console.
Notes from Callum Bennett.