Crawl budget and indexation is the Holistic Radar technical SEO service that gets every page in the topical map crawled and indexed, and keeps everything outside the map from competing for the same crawl. This page defines crawl budget and when it matters, then covers the five causes of wasted crawl, the sitemap rules, canonical and redirect governance, and the index coverage metric measured against the map. It shows a worked audit table, the crawl signals the service monitors, how migrations are handled, and how indexation connects to cost of retrieval. It closes with where the service sits in the Six-Layer Pyramid and the questions buyers ask.
What crawl budget is
Crawl budget is the number of URLs a search engine can and wants to crawl on a site in a given period.
Google describes it as the combination of crawl capacity, how much crawling the server can handle, and crawl demand, how much the engine wants to crawl the site’s URLs. For a site built from a topical map the goal is simple: every map URL is crawled and indexed, and crawl is not spent on URLs the map does not contain.
When crawl budget matters
Crawl budget matters most on sites with hundreds of pages or more, on sites that generate URLs from filters and parameters, and on sites that change often.
A site with a few dozen pages rarely runs out of crawl, but it still needs indexation hygiene. The 400+ page Vital Transportation build is the scale at which wasted crawl starts to delay new pages.
Five causes of wasted crawl
Five causes waste crawl on most sites: duplicate URLs, parameter and filter URLs, thin archive pages, redirect chains, and soft 404s.
| Cause | What happens | Fix |
|---|---|---|
| Duplicate URLs | The same content at two addresses splits crawl and signals | A 301 to the one URL that owns the tuple |
| Parameter and filter URLs | Sort orders and filters create near-infinite URL spaces | Kept out of the index and out of internal links |
| Thin archive pages | Tag, date, and author archives repeat indexed content | Noindex, follow, and removed from the sitemap |
| Redirect chains | Every hop costs a request | Internal links point at final URLs |
| Soft 404s | Empty pages return 200 and get crawled repeatedly | Return a real 404, or a 301 to the right page |
Sitemap rules for crawl budget
The sitemap lists only indexable map pages, with last-modified dates that change only when the content changes.
A sitemap that includes redirected, noindexed, or duplicate URLs teaches the crawler that the sitemap is unreliable. Archives, attachment pages, and author pages are excluded; posts and pages that carry a map tuple are included.
Canonical and redirect governance
Canonical and redirect governance gives every tuple exactly one indexable URL.
Canonical tags point to the owning URL. Merged or retired pages return a 301 to the page that now owns their tuple. A 404 monitor catches old URLs after a migration, and each one found is either redirected to its new owner or left as a real 404. The rules follow Google’s documentation on consolidating duplicate URLs.
Index coverage against the topical map
Index coverage is measured as indexed map URLs divided by total map URLs.
A general indexing report counts every URL the crawler found, including URLs that should not exist. Measuring against the map shows the only number that matters for authority: how much of the planned coverage is actually in the index. The same figure feeds the network integrity component of the Holistic Authority Score.
A worked crawl and indexation audit
A crawl and indexation audit sorts every URL the crawler finds into one of four states and assigns one action to each.
| State | Example | Action |
|---|---|---|
| Map URL, indexed | /services/semantic-seo/ | Keep, monitor |
| Map URL, not indexed | A new node page | Internal links from its hub, sitemap check, request indexing |
| Not in map, duplicate | The same article at two addresses | 301 to the owner |
| Not in map, thin | A tag archive with two posts | Noindex, follow; remove from sitemap |
This site went through the same audit: duplicate Library notes were redirected to the core pages that own their tuples, archives were set to noindex, and articles moved to the map’s /blog/ URLs with 301s from every old address.
Crawl signals the service monitors
The service monitors five crawl signals every month: crawl requests by response code, average response time, indexed map URLs, new 404s, and redirect chains.
Most indexation problems return through site changes rather than decay. A template change that breaks canonicals, or a plugin that starts generating parameter URLs, shows up in these signals before it shows up in rankings.
Crawl budget during migrations
During a migration, every old URL is mapped to one destination before launch, redirects go live with the new site, and the 404 monitor runs from day one.
A migration is also the cheapest moment to fix architecture. Moving URLs straight into the topical map’s structure avoids redirecting the same pages twice.
Crawl budget and cost of retrieval
Crawl budget covers the first of the six stages of cost of retrieval, crawl, and the service keeps it cheap so the later stages get their chance.
Cost of retrieval continues through render, parse, extract, verify, and serve. A page that is not crawled cannot benefit from being well written, which is why the network layer is repaired before the content layer. The full model is in how to lower cost of retrieval.
Crawl budget and indexation in the Six-Layer Pyramid
Crawl budget and indexation sits at layer five of the Six-Layer Pyramid, retrieval and AI confidence, and it is delivered within technical SEO services.
It caps nothing beneath it, but it caps everything the topical map produces: a map page that is never indexed contributes nothing to coverage. The engineering side of the same problem is engineering for cost of retrieval.
Crawl budget and indexation questions
Does a small site need crawl budget work?
A small site rarely has a crawl budget problem, but it still needs indexation hygiene: one URL per tuple, noindexed archives, and an accurate sitemap.
Does noindex save crawl budget?
Noindex keeps a page out of the index but does not stop it being crawled; pages that should not be crawled at all are kept out of internal links and, where appropriate, blocked in robots.txt.
How fast are new map pages indexed?
New map pages are indexed fastest when they are linked from their hub on publication, listed in the sitemap, and served quickly.