Technical SEO · service node · service node

Crawl Budget and Indexation

Every page in your topical map discovered, crawled, and indexed, with nothing outside the map competing for the same crawl.

Crawl budget and indexation is the Holistic Radar technical SEO service that gets every page in the topical map crawled and indexed, and keeps everything outside the map from competing for the same crawl. This page defines crawl budget and when it matters, then covers the five causes of wasted crawl, the sitemap rules, canonical and redirect governance, and the index coverage metric measured against the map. It shows a worked audit table, the crawl signals the service monitors, how migrations are handled, and how indexation connects to cost of retrieval. It closes with where the service sits in the Six-Layer Pyramid and the questions buyers ask.

What crawl budget is

Crawl budget is the number of URLs a search engine can and wants to crawl on a site in a given period.

Google describes it as the combination of crawl capacity, how much crawling the server can handle, and crawl demand, how much the engine wants to crawl the site’s URLs. For a site built from a topical map the goal is simple: every map URL is crawled and indexed, and crawl is not spent on URLs the map does not contain.

When crawl budget matters

Crawl budget matters most on sites with hundreds of pages or more, on sites that generate URLs from filters and parameters, and on sites that change often.

A site with a few dozen pages rarely runs out of crawl, but it still needs indexation hygiene. The 400+ page Vital Transportation build is the scale at which wasted crawl starts to delay new pages.

Five causes of wasted crawl

Five causes waste crawl on most sites: duplicate URLs, parameter and filter URLs, thin archive pages, redirect chains, and soft 404s.

CauseWhat happensFix
Duplicate URLsThe same content at two addresses splits crawl and signalsA 301 to the one URL that owns the tuple
Parameter and filter URLsSort orders and filters create near-infinite URL spacesKept out of the index and out of internal links
Thin archive pagesTag, date, and author archives repeat indexed contentNoindex, follow, and removed from the sitemap
Redirect chainsEvery hop costs a requestInternal links point at final URLs
Soft 404sEmpty pages return 200 and get crawled repeatedlyReturn a real 404, or a 301 to the right page

Sitemap rules for crawl budget

The sitemap lists only indexable map pages, with last-modified dates that change only when the content changes.

A sitemap that includes redirected, noindexed, or duplicate URLs teaches the crawler that the sitemap is unreliable. Archives, attachment pages, and author pages are excluded; posts and pages that carry a map tuple are included.

Canonical and redirect governance

Canonical and redirect governance gives every tuple exactly one indexable URL.

Canonical tags point to the owning URL. Merged or retired pages return a 301 to the page that now owns their tuple. A 404 monitor catches old URLs after a migration, and each one found is either redirected to its new owner or left as a real 404. The rules follow Google’s documentation on consolidating duplicate URLs.

Index coverage against the topical map

Index coverage is measured as indexed map URLs divided by total map URLs.

A general indexing report counts every URL the crawler found, including URLs that should not exist. Measuring against the map shows the only number that matters for authority: how much of the planned coverage is actually in the index. The same figure feeds the network integrity component of the Holistic Authority Score.

A worked crawl and indexation audit

A crawl and indexation audit sorts every URL the crawler finds into one of four states and assigns one action to each.

StateExampleAction
Map URL, indexed/services/semantic-seo/Keep, monitor
Map URL, not indexedA new node pageInternal links from its hub, sitemap check, request indexing
Not in map, duplicateThe same article at two addresses301 to the owner
Not in map, thinA tag archive with two postsNoindex, follow; remove from sitemap

This site went through the same audit: duplicate Library notes were redirected to the core pages that own their tuples, archives were set to noindex, and articles moved to the map’s /blog/ URLs with 301s from every old address.

Crawl signals the service monitors

The service monitors five crawl signals every month: crawl requests by response code, average response time, indexed map URLs, new 404s, and redirect chains.

Most indexation problems return through site changes rather than decay. A template change that breaks canonicals, or a plugin that starts generating parameter URLs, shows up in these signals before it shows up in rankings.

Crawl budget during migrations

During a migration, every old URL is mapped to one destination before launch, redirects go live with the new site, and the 404 monitor runs from day one.

A migration is also the cheapest moment to fix architecture. Moving URLs straight into the topical map’s structure avoids redirecting the same pages twice.

Crawl budget and cost of retrieval

Crawl budget covers the first of the six stages of cost of retrieval, crawl, and the service keeps it cheap so the later stages get their chance.

Cost of retrieval continues through render, parse, extract, verify, and serve. A page that is not crawled cannot benefit from being well written, which is why the network layer is repaired before the content layer. The full model is in how to lower cost of retrieval.

Crawl budget and indexation in the Six-Layer Pyramid

Crawl budget and indexation sits at layer five of the Six-Layer Pyramid, retrieval and AI confidence, and it is delivered within technical SEO services.

It caps nothing beneath it, but it caps everything the topical map produces: a map page that is never indexed contributes nothing to coverage. The engineering side of the same problem is engineering for cost of retrieval.

Crawl budget and indexation questions

Does a small site need crawl budget work?

A small site rarely has a crawl budget problem, but it still needs indexation hygiene: one URL per tuple, noindexed archives, and an accurate sitemap.

Does noindex save crawl budget?

Noindex keeps a page out of the index but does not stop it being crawled; pages that should not be crawled at all are kept out of internal links and, where appropriate, blocked in robots.txt.

How fast are new map pages indexed?

New map pages are indexed fastest when they are linked from their hub on publication, listed in the sitemap, and served quickly.

A useful first step

Find the layer capping your growth.

Start with a six-layer diagnostic and leave with a prioritized next-step sequence.

Get my free Radar ScanFree Radar Scan