• Technical SEO

Setting up sitemap.xml and robots.txt: so the search engine knows what you have

The sitemap.xml and robots.txt settings are two files that decide what a search engine will see on your site. The map lists the addresses you are requesting to bypass; robots.txt tells where not to go and where this map is. We start the work with measurement: we open both addresses and look at the response code, body size, number of records. The most frequent finding in our sample is code 200 with a body of zero bytes. The file seems to exist, the address in it is zero.
See how it works
Price
after a free audit
Guarantee
30 days after project sign-off
What are we doing?
verification of two files, map regeneration, bypass rules, submission to Search Console
after a free audit
Price
30
Guarantee

days after project sign-off

verification of two files, map regeneration, bypass rules, submission to Search Console
What are we doing?
2-10
Term

business days depending on catalog size

13,590
The purest map in its own dimension

addresses and one duplicate for the entire directory

/sitemap.xml returns 200 with a body of zero bytes
The most common sampling defect
free check, 1-2 working days, no accesses
Before the estimate
indexing: the map informs about the pages, the decision is left to the search engine
What we do not promise
What's included

Complete list of work and what you get as a result

  • Regenerate the map with all types of pages and all language versions - the search engine receives the full list, not a part of it.
  • Let's divide the map into an index of submaps, if there are a lot of addresses: the scan report shows which part of the directory was dropped.
  • We will remove service, empty and closed from indexing pages from the map - the bypass is spent on goods, not on the shopping cart.
  • Fix the dates of the last change so that they really update and tell you what to look at after editing the directory.
  • Let's move the card to the standard address, if it is lying on the side: it will be found by the bot, the audit tool and your next contractor.
  • Let's rewrite robots.txt with an explicit Sitemap line — without commented templates from the previous generator.
  • Let's close the service sections, driving each rule through the list of important addresses, so that the category does not disappear along with the office.
  • Let's submit the map to Search Console and Bing Webmaster Tools and wait until it is read without format errors.
  • Let's go through the map with a script after implementation and show two numbers: entries and unique addresses. Matched - the card is clean.
When this service isn't right

What's not included — so there are no surprises at delivery

  • Creation of pages that are not on the site.
  • Indexing guarantee: the map informs about the pages, the decision by the search engine.
  • Working with other people's sites and link directories.
  • Removing duplicates as a separate job — here we just don't drag them into the map.
Who it's for

Situations where this service delivers results

Scenario 1 of 5

The catalog has grown, and the old part of it is visible in the search

The classic situation of a large catalog: new sections are on the site, but the bot reaches them through internal links for months. The map is a direct list of the addresses you are requesting to bypass. In the wholesale underwear catalog, which we measured, there are 13,428 product pages and 161 categories; without a map, the bot would search for the deep positions of such a catalog by itself and for a long time.

We'll review your situation in a free audit

Free sitemap and robots.txt check

The cheapest test possible and often the most effective. We open the standard sitemap and robots.txt URLs and see what they actually return. Access is not required: everything is visible from the outside.

What we measure

  • What /sitemap.xml givesAnswer code, body size, number of entries. We regularly see 200 with a body of zero bytes: formally there is a file, but in fact there is no map.
  • Is the map declared in robots.txtSitemap string. In the sample, there was a commented template, a link to a file with 404 and a complete absence of a directive in the work map.
  • Map completenessHow many addresses is in it and does it match the actual directory. We cross-check with the site's standard search.
  • Entries against unique addressesWe count the two numbers separately. The difference between the two is the doubles that eat the bypass; on other people's measurements, we saw up to 3.9 records per address.
  • Language versions in the mapDid all the announced branches get there? Maps were found where all addresses led to one of the two languages ​​announced in the markings.
  • Index structure and limitsWhether the map is divided into submaps and whether 50,000 addresses and 50 MB per file are not exceeded.
  • Bypass rulesWhat exactly your robots.txt closes and whether it accidentally closes the necessary - along with service sections often close something useful.
  • Card return timeHow many seconds does it take for the file to reach the client. One measured tool store took 34.2 seconds to generate a map — the bot might not wait.

What you get

  • Exact numbers: response code, size, number of entries and unique addresses on your map.
  • A list of what is superfluous in the map and what is missing from it.
  • The ready text of rules for robots.txt for your site.
  • A short conversation, where we go through the findings item by item.

Timeline: 1-2 working days

Why is it free

The check takes us less than an hour, and the result is often unexpected for the owner: the file seems to exist, but the search engine does not get anything from it. Taking money for this would be strange.

What's next

If everything is in order, we will say so, and stop there. If not, let's name a list of works with a deadline and amount; often it is one or two days.

Short form: your contact and site URL

  • Contract, act and 30-day warranty

    Every project gets a written contract: scope, deadlines, amount, acceptance procedure. After delivery — act and invoice, then 30 calendar days of warranty.

  • Sole proprietor & bank transfer

    The contractor is a registered sole proprietor. Payment by invoice with closing documents.

  • Rights & access — yours

    Code, design and materials transfer to you after full payment. Domain, hosting, repository and analytics are registered to you.

  • Client portal instead of email chains

    During the project you get access to a portal: contracts, invoices, acts and project status in one place.

  • European clients

    Among our work — projects for Norway, Bulgaria, Moldova and Spain.

  • Verifiable numbers

    Every case in the portfolio comes with a link to a live site and a technical measurement.

  • Audit first, then pricing

    There is no price list on the site intentionally: the scope of the same work differs multiples between clients.

  • We say "no" when unsure

    If the task isn't ours or the deadline is unrealistic — we tell you upfront.

What affects the price

Why two seemingly identical tasks are priced differently

  • Directory sizeA hundred pages is one map and one passage. Several thousand are an index of submaps, and each one must be checked against the actual catalog. The extreme points of our sample: 123 addresses in the niche store against 13,590 in the wholesale.
  • Who generates the mapThe regular CMS module is configured in a few hours. A self-record generator or external service first requires you to understand where each record comes from and why there are more of them than products.
  • Language versionsOne language - one list. Two or more — all branches should be included in the map, and each branch should be compared with the hreflang markup. Otherwise, half of the site remains unannounced.
  • Robots.txt statusA five-line file is edited in one step. The file, which has been added for years, is analyzed line by line: it usually contains both necessary prohibitions and long-overdue ones.
  • Where to submitVerified by Search Console - same day delivery. Unverified - domain rights are verified first, and the speed here depends on who is holding DNS for you.
Cases

Tasks and results in numbers — all metrics measured by us

wholesale store of underwear, catalog of over 13,000 items

Task
Find out why audit tools don't see the sitemap, even though it exists.
Solution
Checked both standard addresses, read robots.txt, found a working map on its own path and checked its completeness with the directory.
Result
The standard address returns 200 with an empty body, the second one returns a string about the generator being turned off — that is, the two addresses where the crawlers go first return garbage. The working card on a non-standard way contains 13,428 product pages and 161 categories, a total of 13,590 addresses with a single duplicate for the entire volume. It is saved only by the fact that the path is declared in robots.txt. Measured on 07/31/2026.

a niche store of dental materials with a content part

Task
Check the indexing status of a small directory before expanding.
Solution
Bypassed the index file of the map with a list of entries for each submap and a reconciliation with the actual directory.
Result
Correct index of four files: 1 main + 21 categories + 96 products + 5 informational = 123 addresses, zero duplicates. The map here turned out to be the healthiest part of the site — there was something else: not a single analytics counter in the code and not a single block of structured data. Measured on 07/31/2026.

B2B catalog of spare parts for industrial sewing machines

Task
Check why a directory with a full sitemap is poorly represented in search.
Solution
We listed the addresses in the map, checked the language versions, read the robots.txt and what is connected in the code.
Result
The map is clean: 1,992 addresses without any duplicates — 1,848 product cards and 138 categories, hreflang in two languages ​​is correct. But there is no Sitemap directive in robots.txt at all, there is zero structured data on the directory of 1,848 positions, and the analytics system that Google disabled on July 1, 2023 is still connected to the code. Measured on 07/31/2026.

Did not find your case?

Describe how it works on your side — we will tell you whether “Sitemap and robots.txt” fits and what it means in your situation. No brief and no call: one question, one answer.

Process steps

Transparent stages with approval at every step

Total duration:2–10 days

  1. Current state measurement

    1-2 working days

    We open the standard addresses of the map and robots.txt, remove the response code, body size and number of entries, separately count unique addresses. This is the same free check: if everything is in order, there is no need to go further.

  2. A list of what should be in the index

    1 working day

    We compare the catalog with the map: what is missing, what is superfluous. You name sections that deliberately should not be in the search — cabinet, basket, registration, internal search.

  3. Map regeneration

    1-3 working days

    We configure the generator for this list, divide it into sub-maps separately for products, categories and information pages, set the correct change dates. If the card was lying on its side, we transfer it to the standard address.

  4. Bypass rules

    1 working day

    We are rewriting robots.txt: an explicit Sitemap line, service sections are closed, nothing useful was banned. Before filling, run the new file through the list of important addresses.

  5. Viewing and reading the report

    1 working day

    The map is submitted to Search Console and Bing Webmaster Tools. Format errors are visible in hours, not weeks, so this stage is closed quickly or not at all until the card is read cleanly.

  6. Control bypass and transfer

    1-2 working days

    We go through the map with the script again: there are as many entries as there are unique addresses, the return time is fixed. We give the numbers before and after and a short text about what has changed.

Technologies & integrations

What we build on and what it connects to

Stack

  • sitemap.xml is a format read by both Google and Bing; limit: 50,000 addresses and 50 MB per uncompressed file
  • sitemap-index — divides the directory into submaps; restriction: an index cannot refer to another index
  • robots.txt is the only way to tell the bot where the map is; restriction: Disallow prohibits bypassing, not showing in output
  • Search Console is the only place where you can see the actual crawl errors; restrictions: access is provided by the site owner, not us
  • Bing Webmaster Tools is the second point of control; limitations: adds little to Ukrainian traffic
  • in-house OpenCart and WordPress card generators — fast and cheap; restrictions: service and closed pages are drawn into the map
  • own traversal script with the counting of records and unique addresses — catches duplicates that are not visible to the naked eye
  • nginx — a separate map and gzip return rule; limitation: the cache can return the old map after regeneration

Integrations

  • Google Search Console
  • Bing Webmaster Tools
  • Google Merchant Center
  • Cloudflare
A map from a box module vs. a map assembled for your catalog

How this option differs from the alternative

What goes into the cardonly what should be in the index: service, empty and closed pages are not pulled
Doubleswe count records and unique addresses separately, the difference should be zero
Language versionsall declared branches checked with hreflang markup
Size and recoil speedan index of submaps, each given in fractions of a second
Ads in robots.txtan explicit Sitemap string checked against the response code
What we need from you

We can't start without this — best to prepare in advance

  1. Site address - the check starts without any accesses.
  2. Access Search Console to submit a map and see crawl errors with your eyes, not our guesses.
  3. Access to CMS or hosting to replace the file and configure the generator.
  4. A list of sections that deliberately should not be in the index: office, shopping cart, registration, service pages.
  5. One person from your side who will confirm this list - otherwise the bypass rules are written randomly.

If something is missing — let us know, we'll help you gather it or do it as a separate task.

FAQ

Most frequently asked questions — with concrete answers

I have /sitemap.xml open - so everything is ok?

Not necessarily. The most common defect in our sample looks like this: the address returns a code of 200, and the response body is zero bytes. The browser shows an empty page, the owner sees that "the file is there", the search engine does not receive any address. Six sites we measured gave the map in exactly this condition. You need to look at the size of the body and the number of records, and not at the response code.

Do you need a sitemap if everything is already linked on the site?

Needed on a large catalog. The map is a straightforward list of what you think is worth checking out, with dates of change; without it, the bot depends on how deep it will go on the links, and it may not reach the deep pages of the directory for months. For a small site with a good link, Google directly says that a map may not be necessary, and we do not argue with this.

The card is at a non-standard address - is this a problem?

Works if the address is explicitly declared in robots.txt. One of the measured wholesale catalogs was saved precisely by this: the standard addresses gave garbage, and the working map lay on its own path and was registered in robots. But this is a trap for audit tools and for humans, so we transfer the card to a standard address. Another detail of the same level: the addresses inside the map must be on the same host and protocol as the map itself - a map from a subdomain or from an old http bot simply won't count.

How many addresses can be put in one file?

The format limits are 50 thousand addresses and 50 MB in uncompressed form. In practice, it is better to split earlier: in the measured store of the tool, the map was generated with one file in 34.2 seconds, and the bot has a chance not to wait. Our recommendation is a separate subcard for each type of page: products, categories, informational. Then the traversal report shows exactly which part of the directory was dropped, not "something wrong with the map". In the dental supply store, it is done: an index of four files, 1 + 21 + 96 + 5 = 123 addresses.

What is better to hide in robots.txt?

Service sections that have nothing to do in the search: shopping cart, checkout, office, internal search, admin. Be careful with filters and parameters - the prohibition in robots does not remove the page from the index, it only prohibits traversal, and the page may remain in the output without a description. There are other mechanics at work to clean up the issue, and we won't pretend that robots can do it.

The map is submitted, but there are fewer pages in the index - why?

Because the map does not oblige the search engine to index. It's a "here's what we've got" message, not a command, and Google clearly documents it. A product card without a description, without a unique text, and without demand may remain outside the index, no matter how much you submit it. The real number of indexed pages is visible only in your Search Console — 13 thousand addresses in the map do not mean 13 thousand pages in the index.

After cleaning the pagemap, the index became smaller - did we do worse?

No, this is the expected course of events. When duplicates and service addresses disappear from the map, the index counter drops first - simply because the extra stops getting there. The weight is not lost, it stops being scattered between several addresses of the same product. We say this in advance, because without an explanation the fall in the report looks like an accident.

How quickly is it done and what is required of us?

Free verification - 1-2 days and only the website address. More than two days, if you need to regenerate the map and fix the robots; up to ten, if you have to divide the map into submaps, clean it of service addresses and compile it with language versions. From you - access to CMS or hosting, access to Search Console and a list of sections that should not be in the search.

Send the site address - we will start checking the map and robots.txt without access.

In response: the response code, the size and number of addresses in your map, a list of what is redundant and what is missing, and the finished text of the bypass rules. If both files are correct, we will write them — this is also a normal answer.

From measured casesCorrect index of four files: 1 main + 21 categories + 96 products + 5 informational = 123 addresses, zero duplicates

View cases
  • Reply within 2 hours
  • No commitment
  • We work under a contract

There is no price list on the site on purpose: the same work differs several times over between two clients, and a “from” figure explains nothing in that case. First a free audit — we count your pages, duplicates and speed — then we name the sum and the deadline and fix both in the contract.