Setting up sitemap.xml and robots.txt: so the search engine knows what you have
The sitemap.xml and robots.txt settings are two files that decide what a search engine will see on your site. The map lists the addresses you are requesting to bypass; robots.txt tells where not to go and where this map is. We start the work with measurement: we open both addresses and look at the response code, body size, number of records. The most frequent finding in our sample is code 200 with a body of zero bytes. The file seems to exist, the address in it is zero.
verification of two files, map regeneration, bypass rules, submission to Search Console
after a free audit
Price
30
Guarantee
days after project sign-off
verification of two files, map regeneration, bypass rules, submission to Search Console
What are we doing?
2-10
Term
business days depending on catalog size
13,590
The purest map in its own dimension
addresses and one duplicate for the entire directory
/sitemap.xml returns 200 with a body of zero bytes
The most common sampling defect
free check, 1-2 working days, no accesses
Before the estimate
indexing: the map informs about the pages, the decision is left to the search engine
What we do not promise
What's included
Complete list of work and what you get as a result
Regenerate the map with all types of pages and all language versions - the search engine receives the full list, not a part of it.
Let's divide the map into an index of submaps, if there are a lot of addresses: the scan report shows which part of the directory was dropped.
We will remove service, empty and closed from indexing pages from the map - the bypass is spent on goods, not on the shopping cart.
Fix the dates of the last change so that they really update and tell you what to look at after editing the directory.
Let's move the card to the standard address, if it is lying on the side: it will be found by the bot, the audit tool and your next contractor.
Let's rewrite robots.txt with an explicit Sitemap line — without commented templates from the previous generator.
Let's close the service sections, driving each rule through the list of important addresses, so that the category does not disappear along with the office.
Let's submit the map to Search Console and Bing Webmaster Tools and wait until it is read without format errors.
Let's go through the map with a script after implementation and show two numbers: entries and unique addresses. Matched - the card is clean.
When this service isn't right
What's not included — so there are no surprises at delivery
Creation of pages that are not on the site.
Indexing guarantee: the map informs about the pages, the decision by the search engine.
Working with other people's sites and link directories.
Removing duplicates as a separate job — here we just don't drag them into the map.
Who it's for
Situations where this service delivers results
Scenario 1 of 5
The catalog has grown, and the old part of it is visible in the search
The classic situation of a large catalog: new sections are on the site, but the bot reaches them through internal links for months. The map is a direct list of the addresses you are requesting to bypass. In the wholesale underwear catalog, which we measured, there are 13,428 product pages and 161 categories; without a map, the bot would search for the deep positions of such a catalog by itself and for a long time.
We'll review your situation in a free audit
Scenario 2 of 5
The audit tool says "sitemap not found" even though you can see it
It means that the map is not where it is being sought. In the same wholesale directory, both standard addresses returned garbage—one an empty body, the other a line about the generator being turned off—while the work card lived on its own path. The only thing that saved her was that the path was declared in robots.txt.
Scenario 3 of 5
There are two languages on the site, but one was included in the map
Half of the site is simply not announced. We have found such maps more than once: in the marking there are two language branches, in the map 100% of addresses lead to one. Here, hreflang is also applied — and while the map is incomplete, there is nothing to play with.
Scenario 4 of 5
You have closed a section in robots.txt and it is still in issue
Disallow disallows bypassing, not showing. The page may remain in the results without a description, because the bot no longer has the right to enter and read what is inside. This is the first thing we explain when asked to "just close filters in robots".
Scenario 5 of 5
The site has just been taken over - moved, changed the structure, changed the CMS
After such work, the map almost always lags behind the site: it keeps old addresses and does not know new ones. Along with it, robots.txt lags behind, in which prohibitions from the previous structure remain.
Free sitemap and robots.txt check
The cheapest test possible and often the most effective. We open the standard sitemap and robots.txt URLs and see what they actually return. Access is not required: everything is visible from the outside.
What we measure
What /sitemap.xml givesAnswer code, body size, number of entries. We regularly see 200 with a body of zero bytes: formally there is a file, but in fact there is no map.
Is the map declared in robots.txtSitemap string. In the sample, there was a commented template, a link to a file with 404 and a complete absence of a directive in the work map.
Map completenessHow many addresses is in it and does it match the actual directory. We cross-check with the site's standard search.
Entries against unique addressesWe count the two numbers separately. The difference between the two is the doubles that eat the bypass; on other people's measurements, we saw up to 3.9 records per address.
Language versions in the mapDid all the announced branches get there? Maps were found where all addresses led to one of the two languages announced in the markings.
Index structure and limitsWhether the map is divided into submaps and whether 50,000 addresses and 50 MB per file are not exceeded.
Bypass rulesWhat exactly your robots.txt closes and whether it accidentally closes the necessary - along with service sections often close something useful.
Card return timeHow many seconds does it take for the file to reach the client. One measured tool store took 34.2 seconds to generate a map — the bot might not wait.
What you get
Exact numbers: response code, size, number of entries and unique addresses on your map.
A list of what is superfluous in the map and what is missing from it.
The ready text of rules for robots.txt for your site.
A short conversation, where we go through the findings item by item.
Timeline: 1-2 working days
Why is it free
The check takes us less than an hour, and the result is often unexpected for the owner: the file seems to exist, but the search engine does not get anything from it. Taking money for this would be strange.
What's next
If everything is in order, we will say so, and stop there. If not, let's name a list of works with a deadline and amount; often it is one or two days.
Short form: your contact and site URL
Contract, act and 30-day warranty
Every project gets a written contract: scope, deadlines, amount, acceptance procedure. After delivery — act and invoice, then 30 calendar days of warranty.
Sole proprietor & bank transfer
The contractor is a registered sole proprietor. Payment by invoice with closing documents.
Rights & access — yours
Code, design and materials transfer to you after full payment. Domain, hosting, repository and analytics are registered to you.
Client portal instead of email chains
During the project you get access to a portal: contracts, invoices, acts and project status in one place.
European clients
Among our work — projects for Norway, Bulgaria, Moldova and Spain.
Verifiable numbers
Every case in the portfolio comes with a link to a live site and a technical measurement.
Audit first, then pricing
There is no price list on the site intentionally: the scope of the same work differs multiples between clients.
We say "no" when unsure
If the task isn't ours or the deadline is unrealistic — we tell you upfront.
What affects the price
Why two seemingly identical tasks are priced differently
Directory sizeA hundred pages is one map and one passage. Several thousand are an index of submaps, and each one must be checked against the actual catalog. The extreme points of our sample: 123 addresses in the niche store against 13,590 in the wholesale.
Who generates the mapThe regular CMS module is configured in a few hours. A self-record generator or external service first requires you to understand where each record comes from and why there are more of them than products.
Language versionsOne language - one list. Two or more — all branches should be included in the map, and each branch should be compared with the hreflang markup. Otherwise, half of the site remains unannounced.
Robots.txt statusA five-line file is edited in one step. The file, which has been added for years, is analyzed line by line: it usually contains both necessary prohibitions and long-overdue ones.
Where to submitVerified by Search Console - same day delivery. Unverified - domain rights are verified first, and the speed here depends on who is holding DNS for you.
Cases
Tasks and results in numbers — all metrics measured by us
wholesale store of underwear, catalog of over 13,000 items
Task
Find out why audit tools don't see the sitemap, even though it exists.
Solution
Checked both standard addresses, read robots.txt, found a working map on its own path and checked its completeness with the directory.
Result
The standard address returns 200 with an empty body, the second one returns a string about the generator being turned off — that is, the two addresses where the crawlers go first return garbage. The working card on a non-standard way contains 13,428 product pages and 161 categories, a total of 13,590 addresses with a single duplicate for the entire volume. It is saved only by the fact that the path is declared in robots.txt. Measured on 07/31/2026.
a niche store of dental materials with a content part
Task
Check the indexing status of a small directory before expanding.
Solution
Bypassed the index file of the map with a list of entries for each submap and a reconciliation with the actual directory.
Result
Correct index of four files: 1 main + 21 categories + 96 products + 5 informational = 123 addresses, zero duplicates. The map here turned out to be the healthiest part of the site — there was something else: not a single analytics counter in the code and not a single block of structured data. Measured on 07/31/2026.
B2B catalog of spare parts for industrial sewing machines
Task
Check why a directory with a full sitemap is poorly represented in search.
Solution
We listed the addresses in the map, checked the language versions, read the robots.txt and what is connected in the code.
Result
The map is clean: 1,992 addresses without any duplicates — 1,848 product cards and 138 categories, hreflang in two languages is correct. But there is no Sitemap directive in robots.txt at all, there is zero structured data on the directory of 1,848 positions, and the analytics system that Google disabled on July 1, 2023 is still connected to the code. Measured on 07/31/2026.
Did not find your case?
Describe how it works on your side — we will tell you whether “Sitemap and robots.txt” fits and what it means in your situation. No brief and no call: one question, one answer.
Process steps
Transparent stages with approval at every step
Total duration:2–10 days
1
Current state measurement
1-2 working days
We open the standard addresses of the map and robots.txt, remove the response code, body size and number of entries, separately count unique addresses. This is the same free check: if everything is in order, there is no need to go further.
2
A list of what should be in the index
1 working day
We compare the catalog with the map: what is missing, what is superfluous. You name sections that deliberately should not be in the search — cabinet, basket, registration, internal search.
3
Map regeneration
1-3 working days
We configure the generator for this list, divide it into sub-maps separately for products, categories and information pages, set the correct change dates. If the card was lying on its side, we transfer it to the standard address.
4
Bypass rules
1 working day
We are rewriting robots.txt: an explicit Sitemap line, service sections are closed, nothing useful was banned. Before filling, run the new file through the list of important addresses.
5
Viewing and reading the report
1 working day
The map is submitted to Search Console and Bing Webmaster Tools. Format errors are visible in hours, not weeks, so this stage is closed quickly or not at all until the card is read cleanly.
6
Control bypass and transfer
1-2 working days
We go through the map with the script again: there are as many entries as there are unique addresses, the return time is fixed. We give the numbers before and after and a short text about what has changed.
Technologies & integrations
What we build on and what it connects to
Stack
sitemap.xml is a format read by both Google and Bing; limit: 50,000 addresses and 50 MB per uncompressed file
sitemap-index — divides the directory into submaps; restriction: an index cannot refer to another index
robots.txt is the only way to tell the bot where the map is; restriction: Disallow prohibits bypassing, not showing in output
Search Console is the only place where you can see the actual crawl errors; restrictions: access is provided by the site owner, not us
Bing Webmaster Tools is the second point of control; limitations: adds little to Ukrainian traffic
in-house OpenCart and WordPress card generators — fast and cheap; restrictions: service and closed pages are drawn into the map
own traversal script with the counting of records and unique addresses — catches duplicates that are not visible to the naked eye
nginx — a separate map and gzip return rule; limitation: the cache can return the old map after regeneration
Integrations
Google Search Console
Bing Webmaster Tools
Google Merchant Center
Cloudflare
A map from a box module vs. a map assembled for your catalog
How this option differs from the alternative
Map "how the module generated"Our approach
What goes into the cardeverything the generator has reached, including the recycle bin and filter pagesonly what should be in the index: service, empty and closed pages are not pulled
Doublesthe same product is mapped under each category pathwe count records and unique addresses separately, the difference should be zero
Language versionsoften one language out of two — and no one will know about itall declared branches checked with hreflang markup
Size and recoil speedone file for the entire directory; in one measured store it generated 34.2 secondsan index of submaps, each given in fractions of a second
Ads in robots.txtboilerplate commented string or 404 linkan explicit Sitemap string checked against the response code
What we need from you
We can't start without this — best to prepare in advance
1Site address - the check starts without any accesses.
2Access Search Console to submit a map and see crawl errors with your eyes, not our guesses.
3Access to CMS or hosting to replace the file and configure the generator.
4A list of sections that deliberately should not be in the index: office, shopping cart, registration, service pages.
5One person from your side who will confirm this list - otherwise the bypass rules are written randomly.
If something is missing — let us know, we'll help you gather it or do it as a separate task.
FAQ
Most frequently asked questions — with concrete answers
I have /sitemap.xml open - so everything is ok?
Not necessarily. The most common defect in our sample looks like this: the address returns a code of 200, and the response body is zero bytes. The browser shows an empty page, the owner sees that "the file is there", the search engine does not receive any address. Six sites we measured gave the map in exactly this condition. You need to look at the size of the body and the number of records, and not at the response code.
Do you need a sitemap if everything is already linked on the site?
Needed on a large catalog. The map is a straightforward list of what you think is worth checking out, with dates of change; without it, the bot depends on how deep it will go on the links, and it may not reach the deep pages of the directory for months. For a small site with a good link, Google directly says that a map may not be necessary, and we do not argue with this.
The card is at a non-standard address - is this a problem?
Works if the address is explicitly declared in robots.txt. One of the measured wholesale catalogs was saved precisely by this: the standard addresses gave garbage, and the working map lay on its own path and was registered in robots. But this is a trap for audit tools and for humans, so we transfer the card to a standard address. Another detail of the same level: the addresses inside the map must be on the same host and protocol as the map itself - a map from a subdomain or from an old http bot simply won't count.
How many addresses can be put in one file?
The format limits are 50 thousand addresses and 50 MB in uncompressed form. In practice, it is better to split earlier: in the measured store of the tool, the map was generated with one file in 34.2 seconds, and the bot has a chance not to wait. Our recommendation is a separate subcard for each type of page: products, categories, informational. Then the traversal report shows exactly which part of the directory was dropped, not "something wrong with the map". In the dental supply store, it is done: an index of four files, 1 + 21 + 96 + 5 = 123 addresses.
What is better to hide in robots.txt?
Service sections that have nothing to do in the search: shopping cart, checkout, office, internal search, admin. Be careful with filters and parameters - the prohibition in robots does not remove the page from the index, it only prohibits traversal, and the page may remain in the output without a description. There are other mechanics at work to clean up the issue, and we won't pretend that robots can do it.
The map is submitted, but there are fewer pages in the index - why?
Because the map does not oblige the search engine to index. It's a "here's what we've got" message, not a command, and Google clearly documents it. A product card without a description, without a unique text, and without demand may remain outside the index, no matter how much you submit it. The real number of indexed pages is visible only in your Search Console — 13 thousand addresses in the map do not mean 13 thousand pages in the index.
After cleaning the pagemap, the index became smaller - did we do worse?
No, this is the expected course of events. When duplicates and service addresses disappear from the map, the index counter drops first - simply because the extra stops getting there. The weight is not lost, it stops being scattered between several addresses of the same product. We say this in advance, because without an explanation the fall in the report looks like an accident.
How quickly is it done and what is required of us?
Free verification - 1-2 days and only the website address. More than two days, if you need to regenerate the map and fix the robots; up to ten, if you have to divide the map into submaps, clean it of service addresses and compile it with language versions. From you - access to CMS or hosting, access to Search Console and a list of sections that should not be in the search.
Send the site address - we will start checking the map and robots.txt without access.
In response: the response code, the size and number of addresses in your map, a list of what is redundant and what is missing, and the finished text of the bypass rules. If both files are correct, we will write them — this is also a normal answer.
From measured casesCorrect index of four files: 1 main + 21 categories + 96 products + 5 informational = 123 addresses, zero duplicates
There is no price list on the site on purpose: the same work differs several times over between two clients, and a “from” figure explains nothing in that case. First a free audit — we count your pages, duplicates and speed — then we name the sum and the deadline and fix both in the contract.