• Technical SEO · Search Console

Crawling budget: why the bot does not reach half of the directory

Optimizing the crawling budget is working with what the bot spends on a limited number of requests to your site. We take scan statistics, calculate the share of traffic that goes to repetitions and service addresses, remove them and accelerate the return of pages that the bot opens most often. The size of the directory here is secondary: in the tool's measured store, the sitemap returned 25,537 records for 6,857 real addresses — almost three-quarters of the bot's requests went to re-circulating already known ones.

See how it works
Price
after a free audit
Guarantee
30 days after project sign-off
What are we doing?
scan statistics, bypass sweep, recoil rate, control slice
after a free audit
Price
30
Guarantee

days after project sign-off

scan statistics, bypass sweep, recoil rate, control slice
What are we doing?
10-25
The term of the main works

working days

59,189
The largest catalog in its own dimension

addresses in the site map, of which 55,321 are commercial

25,537
How many repetitions eat up

card records for 6,857 addresses — 3.9 repetitions per product

free bypass estimate, 3-4 business days
Before the estimate
the number of pages in the index — the bypass does not determine it
What we do not promise

A free estimate of how much traffic your site is spending

The crawling budget is not an abstraction, but a specific number of requests that the bot makes to your site. The only question is what they go for: for new goods or for the tenth repetition of the same card. The assessment of this composition costs nothing.

What we measure

  • How unique is the site mapEntries against unique addresses. Every extra repetition is a bot request wasted.
  • Recoil speed under the botHow much time the server spends on one page. The slower, the fewer pages the bot manages per session.
  • Site map generation timeWe measured the store where the card was issued in 34.2 seconds. The bot does not have to wait.
  • Service addresses are bypassedInternal lookups, sorting, comparisons, sessions — anything that shouldn't be in the index but is traversable.
  • Language branchesDoes the traversal not double due to languages ​​and do both versions need the same in the index.
  • Scan statisticsSearch Console data: how many pages per day, what is the average response time, how many errors. This is the single source of truth about bot behavior.

What you get

  • Estimation of the share of bypass going to repeaters and service addresses.
  • A list of what to remove from the bypass, with an explanation of the mechanics for each item.
  • Average response times compared to what we've seen on other sites in the same class.
  • Conversation for 30-40 minutes on the document.

Timeline: 3-4 working days

Why is it free

Because without access to scan statistics, any estimate of the bypass is a guess, and we can make the first cut from the outside. It's more honest than selling work blindly.

What's next

Next is a list of works with the amount and term. Some of the points overlap with cleaning duplicates: if it is more profitable for you to do one job instead of two, we will say so.

Short form: your contact and site URL

  • Contract, act and 30-day warranty

    Every project gets a written contract: scope, deadlines, amount, acceptance procedure. After delivery — act and invoice, then 30 calendar days of warranty.

  • Sole proprietor & bank transfer

    The contractor is a registered sole proprietor. Payment by invoice with closing documents.

  • Rights & access — yours

    Code, design and materials transfer to you after full payment. Domain, hosting, repository and analytics are registered to you.

  • Client portal instead of email chains

    During the project you get access to a portal: contracts, invoices, acts and project status in one place.

  • European clients

    Among our work — projects for Norway, Bulgaria, Moldova and Spain.

  • Verifiable numbers

    Every case in the portfolio comes with a link to a live site and a technical measurement.

  • Audit first, then pricing

    There is no price list on the site intentionally: the scope of the same work differs multiples between clients.

  • We say "no" when unsure

    If the task isn't ours or the deadline is unrealistic — we tell you upfront.

What affects the price

Why two seemingly identical tasks are priced differently

  • How many addresses does the bot know?Prime factor. The inventory of the directory for 590 addresses and the directory for 59,189 are different jobs not in percentages, but in weeks: submaps have to be bypassed actually, page by page.
  • Share of repetitionsRepetitions are not analyzed piece by piece - they fall into patterns, and one rule is written on the pattern. But the template must be found first. In the measured tool store, one address accounted for 3.9 entries in the sitemap.
  • Are there any server logs?With the logs, you can see where the bot actually goes, including addresses that are not in the sitemap or in the links. Without them, Search Console remains: aggregated numbers without breakdown by address, and the path to the same conclusions is longer.
  • Who manages the serverWhen the cache, robots.txt and sitemap generator are managed by us, it is calendar fast. When everything goes through your admin or restricted hosting, the time to approve each edit is added.
  • The number of language branchesTwo languages ​​are not twice as much work, but plus a separate layer of reconciliation: are both branches in the site map and does the language of the addresses match the language that the site displays by default.
  • Site map statusA map assembled on the fly is a separate line of work. We measured the store, where it took 34.2 seconds to generate it: there the issue is no longer in cleaning records, but in switching to static submaps.
Who it's for

Situations where this service delivers results

Scenario 1 of 4

A catalog of thousands of addresses, which is regularly updated

The product appeared on the site on Monday, in the search - after three weeks or never. In a catalog of this size, the bot does not have time to go through everything in one pass, and the queue is not built according to your priority. The job is to make room in this queue: remove from it what should not be in the index, and leave new cards.

We'll review your situation in a free audit
What's included

Complete list of work and what you get as a result

  • Analysis of Search Console crawl statistics for the available period: pages per day, average response time, distribution of response codes - the point from which the rest is measured
  • Calculation of the actual composition of the round: how many cards, categories, service addresses and repetitions are in it. Next, you see the share, not the feeling
  • Closing service sections from traversal — internal search, sorting, comparison, trash, session IDs: bot requests are no longer wasted on them
  • Sitemap cleanup: one address - one entry so the bot gets a list of products instead of the same product under four category paths
  • Reducing the time of collecting a site map or switching to static submaps — the bot does not have to wait for it to be generated
  • Acceleration of pages that the bot opens most often: cache, requests to the database, nginx settings. This directly adds pages per session
  • True last modified dates so the bot sees what's actually updated and doesn't go through the entire directory blindly
  • View internal linking: deep directory pages are pulled closer to the entrance, because there should be fewer transitions to them
  • Control slice of scan statistics 30 days after implementation - with the same metrics that were taken at the start
  • A written explanation of each change so your developer can continue to support it without us
When this service isn't right

What's not included — so there are no surprises at delivery

  • Guarantee of the number of pages in the index: traversal does not determine it
  • Acceleration of indexing of individual pages "on demand"
  • Creating content for pages that are left out of the index due to lack of value
  • Working with external links
  • Rewriting the driver, if the address repetitions are embedded in the platform itself, is a separate project with a separate estimate
Process steps

Transparent stages with approval at every step

Total duration:10–25 days

  1. Scan statistics and logs

    2-4 working days

    We extract the Search Console report for the available period and, if access is available, the server logs. We look at the volume of traffic per day, the average response time, the distribution of codes. Right away, it becomes clear whether the problem is in the bypass budget at all.

  2. An inventory of what the bot bypasses

    2-5 working days

    Traversing the site map with a count of records and unique addresses, running the structure with a crawler, breaking it down into cards, categories, service addresses and language branches. The output is a table with the share of each type in the general round.

  3. Bypass cleaning

    3-7 working days

    Service sections are closed, repetitions are removed from the site map, canonical addresses are placed where one page is accessible in several ways. We explain each decision in writing: bypass closure and de-indexing are different mechanisms, and confusion between them is costly.

  4. Site map and response time

    2-5 working days

    The map is converted to submaps or static generation, the dates of the last change begin to correspond to reality. At the same time, we take care of the return of those pages that the bot opens most often - usually these are categories and the first pages of the pagination.

  5. Linking and transmission

    1-4 working days

    Deep directory pages are pulled closer to the login, all changes are recorded in a document for your developer. Right here, we agree on the date of the control cut and who will watch the report after us.

  6. Control section

    30 days after implementation, outside the work period

    Repeat scan statistics using the same metrics. There is no point in looking earlier: Search Console data comes with a delay, and the bot's behavior does not change in a week.

Did not find your case?

Describe how it works on your side — we will tell you whether “Optimization of the crawling budget” fits and what it means in your situation. No brief and no call: one question, one answer.

Cases

Tasks and results in numbers — all metrics measured by us

online store of professional tools, catalog of over 6,000 items

Task
Understand why new positions do not appear in the search for a long time, although the site map is filled.
Solution
We went around the site map with a count of records and unique addresses, measured the time of its collection, checked the language composition of the addresses with the language that the site gives by default.
Result
25,537 records for 6,857 unique addresses: 6,486 product cards and 361 categories — each product is present in the card an average of 3.9 times. The map itself is assembled on the fly for 34.2 seconds, lies on a non-standard path and is only found through a directive in robots.txt. All addresses in it are in one language branch, although by default the site gives another. Measured on 07/31/2026.

a brand of natural spices with its own online store

Task
Check why a part of the catalog does not get issued with a small volume of goods.
Solution
Counting of entries and unique addresses in the site map, comparison of declared language markings with the actual composition of addresses, inventory of analytics counters.
Result
3,971 records for 590 unique addresses with a catalog of about 580 products. At this volume, the bot bypasses everything, and the bottleneck was not the budget, but the direction of the bypass: all 590 addresses lead to one language version, although the document is declared in a different language. The first response is 428.4 ms with 264.1 KB of markup. Measured on 07/31/2026.

online store of garden equipment and spare parts, catalog of over 50,000 items

Task
Find out whether the bot has time to bypass a directory of this size and where it slows down.
Solution
We actually counted the addresses in the submaps, page by page, measured the time of the first response and the weight of the markup, checked how the map is declared in robots.txt.
Result
55,321 commodity addresses are spread over 19 sub-cards — 3,000 in each, except for the last of the 1,321 — plus 3,868 categorical: at least 59,189 addresses. First response 840.5ms, markup weight 454.5KB, PHP 7.3.33 not supported since December 2021. The map is declared in robots.txt, but lies on a non-standard path. Measured on 07/31/2026.
Technologies & integrations

What we build on and what it connects to

Stack

  • Search Console, crawl statistics — a single source of bot behavior; data comes with a delay
  • Server logs — an exact trace of each visit; there is usually no access to them on shared hosting
  • Screaming Frog — bypassing the structure with a crawler; it won't find a page without any incoming link either
  • robots.txt — directly removes the load from the crawl, but does not remove from the index what has already been there
  • Canonical addresses are on the contrary: they paste duplicates in the index, but they do not save on traversal
  • sitemap-index from submaps — so that the map is not assembled on the fly; working benchmark — 3,000 addresses per subcard
  • Last-Modified and ETag — signal "has not changed"; only works as long as the date is valid, not the current time
  • nginx and Cloudflare — cache and response time; the long tail of the directory still bypasses the cache

Integrations

  • Google Search Console
  • Bing Webmaster Tools
  • Cloudflare
  • GA4
Work with bypass composition versus manually sending addresses for bypass

How this option differs from the alternative

What is changingcomposition of the bypass: the bot stops walking on repetitions and service addresses
Scalethe effect is greater, the more addresses in the directory
Durationit lasts until a new layer of repetitions grows
What can be seen in the reportpages per day, average response time, distribution of codes — before and after
What we need from you

We can't start without this — best to prepare in advance

  1. Access to Search Console: Without crawl statistics, this job is half blind, and we will say so.
  2. Access to server logs, if available at all, is the most accurate source of where the bot goes.
  3. A list of sections that are deliberately not needed in the index: selections for advertising, service filters, archives.
  4. Understanding how often the catalog is updated - daily, weekly or once a season. It depends on what to chase.
  5. A development person who can change robots.txt, sitemap generator and bounce settings.
  6. Decisions about language branches, if there are two of them: both remain in the index or one.

If something is missing — let us know, we'll help you gather it or do it as a separate task.

FAQ

Most frequently asked questions — with concrete answers

I have 500 products - does the crawl budget apply to me?

Practically not, and it would be dishonest to sell you this work. In the directory of several hundred addresses, the bot bypasses everything, and even with a margin - the bottleneck is somewhere else. The work begins to pay off at thousands of addresses and especially where the catalog is frequently replenished. The spice shop we measured is just such a case: 590 unique addresses, and the problem was not in the budget of the detour, but in the fact that they all led to one language branch out of the two announced.

How to understand that the bypass is spent in vain?

The fastest check takes five minutes: open the site map and compare the number of records with the number of unique addresses. If the first is multiple times the second, the bot regularly walks on the same path. In the professional tool store, we counted 25,537 records for 6,857 addresses — 3.9 repetitions per item. The second check is more difficult: in the scan statistics report, the average response time is increasing, while the number of pages per day is decreasing. Then it's not a matter of structure, but of speed.

Does site speed affect crawling?

Direct. The longer the server takes to render the page, the less their bot can do per session. For a garden equipment store with 59,189 addresses, the first response takes 840.5ms, and at this volume, that's thousands of uncirculated pages per pass. A special case is the site map itself: when it takes 34.2 seconds, the question is no longer the speed of the pages, but whether the bot will wait for the list at all. Fair margin: the first answer is the server part, it does not describe what a live visitor sees from the phone, it adds images, scripts and network.

Is it enough to close the unnecessary in robots.txt?

To save a detour — yes, it works directly and immediately. But closing from the bypass does not remove the page from the index, if it has already been there: the bot stops opening it, and it remains in the output - without a description, because there is no where to read it. De-indexing requires a different mechanism, and we mention in each point of the report which one was used and why.

What to do with language versions?

Deciding whether both are needed in the index first is a business decision, not ours. Each branch multiplies the list of addresses, and with it, the detour. In the two measured stores, the sitemap contained only one branch, and in both cases it was not the one that the site gives by default, meaning that the detour bypassed the main version. The other side of the issue: the language branch is also content that someone has to maintain all the time, and this expense never ends.

How long until the visible effect?

Crawl statistics start to move within a few weeks - no sooner, Search Console data comes with a delay. The practical effect, that is, the faster appearance of new products in the search, becomes noticeable in the next catalog update cycle. We make a control section 30 days after implementation and compare the same metrics that were taken at the start.

Will the number of pages in the index drop after cleaning?

It may fall, and this is the expected course of events. Duplicates and business addresses are removed from the index, so the number in the report decreases - at the expense of what never brought visits. We are talking about this before the work starts, precisely because the drop of the counter without an explanation is read as a deterioration. And separately: the number of addresses in the site map is never equal to the number of indexed pages - this is how Google describes it.

Does more crawl mean more traffic?

No, and this is an important boundary. A faster crawl means that the page becomes available in search sooner - it does not create demand for it. A card without a description, without distinction from the neighboring one and without requests in the niche can get into the index in a day and not receive a single visit. Therefore, in our metrics, there is a period from publication to appearance in the index, and not a promise of an increase in visits.

Submit site address and access to crawl statistics.

In response, the share of the detour that goes to repeaters and service addresses, the time of collecting your sitemap, and a list of things that should be removed. If the catalog is small and you don't mind going around, let's say so in the first email.

From measured cases25,537 records for 6,857 unique addresses: 6,486 product cards and 361 categories

View cases
  • Reply within 2 hours
  • No commitment
  • We work under a contract

There is no price list on the site on purpose: the same work differs several times over between two clients, and a “from” figure explains nothing in that case. First a free audit — we count your pages, duplicates and speed — then we name the sum and the deadline and fix both in the contract.