Semantic core for an online store: which categories should exist and which are missing
The semantic core for an online store is not a query file, but a list of pages that should exist in the directory. Each group of requests has its own category, subcategory or landing page, and it is this list that you receive at the exit. We start the work not by collecting, but by comparing: what you already have, what people are looking for, and where these two lists diverge. The most expensive find is a request for which the product is in stock, but there is no page on the website.
working days depending on the width of the assortment
free, 3-5 working days
Reconciliation of structure with demand
the draft structure and technical specifications for changes, not a table of frequencies
What's out
Process steps
Transparent stages with approval at every step
Total duration:15–40 days
1
Reconciliation of structure with demand
2-4 working days
We remove the current category tree, count the directory using two methods, look at the internal search and Search Console at the top level. This is the same free part: its result becomes the basis for evaluating the entire work.
2
Collection of inquiries by product range
4-8 working days
We do not proceed from the alphabet, but from the export of the price: each direction of the assortment is closed separately, together with the brands, compatibility and characteristics by which the product is sought. We immediately filter out the irrelevant.
3
Clustering
3-7 working days
We group requests according to the way the search engine sees them: if the output for two requests mostly coincides, they are served by one page. At the boundaries of groups, decisions are made by hand — the automation is most often wrong there.
4
Superimposing clusters on existing pages
2-5 working days
Each cluster receives a status: there is a page, there are two pages, there is no page. Empty sections and categories compete for the same group of requests.
5
Project structure and prioritization
2-8 working days
Category tree with nesting levels, names and addresses. Separately, we divide what becomes a category, what remains a filter, and what is taken to the boarding area. The work queue is built at the intersection of demand and margin.
6
Technical specifications for changes and a map of relinking
2-8 working days
Tasks for a developer or content manager: what to create, what to combine, what to remove, which addresses change and which transitions from the old ones are needed. Plus a link diagram between new and existing sections.
Technologies & integrations
What we build on and what it connects to
Stack
Search Console — queries for which you are already shown; there is no new demand there by definition, and the data comes with a delay
Internal site search is the most accurate source of what is being searched for and not found on your site; silent where it is off
Clustering by intersection of output — groups queries as they are seen by the search engine; output changes, so large clusters are checked again
Screaming Frog — removes structure, nesting and empty sections; a complete bypass of tens of thousands of addresses takes hours, so we do it in branches
The site map as a source of the list of nodes — in the spare parts store, we removed 3,868 category addresses from the submaps; works as long as the card does not lie
GA4 — shows which sections give orders; it's a trend, not a cash register
Export the price is the source of the truth about what you really have; without it, demand without a category is indistinguishable from demand without a product
Integrations
Google Search Console
GA4
internal site search
Google Merchant Center
What's included
Complete list of work and what you get as a result
Collection of requests in all areas of the assortment with cleaning from irrelevant — to continue working with demand, and not with parser garbage.
Clustering: Groups of requests, each served by a single page. The number of clusters shows the real amount of work ahead.
Superimposing clusters on existing pages: under which there is already an address, under which there are two, and under which there is none.
List of categories with no demand - sections whose names no one searches for. They don't hurt, but they don't work either, and you decide on each one.
Project directory structure: which categories to create, which to combine, which to remove and at what nesting level each is.
The rules for naming categories and cards that scale to the entire catalog — the next hundred items are created without us.
Prioritizing at the intersection of demand and marginality: what makes money first, not what is most in demand.
Terms of reference for changes to the structure for a developer or content manager - with a division into what affects addresses and what does not affect them.
A link map between new and existing sections so that new pages are not left without any internal links.
When this service isn't right
What's not included — so there are no surprises at delivery
Writing texts for new categories is a separate job and a separate budget
Implementing CMS changes if your team does
Promises of items based on collected requests
Work with advertising campaigns
Assortment changes: we show the demand, the decision to buy or not is up to you
Who it's for
Situations where this service delivers results
Scenario 1 of 5
The catalog grew from the supplier's price
The sections are named as the supplier named them: Consumables, Group 4, Other. People are searching for the same products with other words, and there is no common root between the title of the section and the query. This is the cheapest discrepancy to fix: the product is there, the page is there, only the name and address do not match.
We'll review your situation in a free audit
Scenario 2 of 5
The product is in stock, but there is no page for it
The most expensive type of break. The query group exists, the query is generated, the items are within some broad category - and no address to show for this query. This is the intersection of the price download with internal search data: people ask with words inside the site exactly what they did not find in the menu.
Scenario 3 of 5
Two categories share one request
Sections with almost the same content, created in different years by different people. The searcher chooses one of them himself, and not always the one with the best assortment; part of the links and behavioral signals goes to the second. The solution here is not a technical one, but an assortment one: which one we leave, which one we transfer, where the address of the one that disappears leads to.
Scenario 4 of 5
Catalog of tens of thousands of items
In the store of garden equipment and spare parts measured by us, there are 55,321 product addresses and 3,868 category nodes. At this volume, the structure is no longer a question of convenience, but a question of whether it is even possible to get to the right spare part. The semantics here correspond to exactly which nodes should be and at what depth.
Scenario 5 of 5
The catalog is being rebuilt or moved
The moment when this work gives the most: the structure has not yet been filled in the CMS, the new addresses have not yet been indexed, transitions from old pages are not yet possible. If the move is planned, the semantics should be done before it.
Semantics for a structure versus a core file
How this option differs from the alternative
kernel file for thousands of requestsOur approach
What exactly is given at the exitquery strings grouped by frequencya list of pages that should exist, with a solution for each
How to measure the amount of workby the number of requests in the fileby the number of clusters - a thousand requests can be reduced to forty pages
What to do nextthe client himself decides where to send groups of requestsTOR for structure changes with a list of addresses and transitions
Empty and duplicate partitionsare not considered - they are not among the requestsare presented in a separate list with a decision: to fill, reduce or remove
Filters and combinations of characteristicseverything in demand is offered as a separate pagedivided before implementation: category, filter without address or landing
Order of workwith decreasing frequencyat the intersection of demand and direct marginality
Free comparison of the structure of the catalog with the demand
Semantics only make sense when a structure grows out of it. Therefore, the first step is not to collect requests, but to compare: what sections you have, what demand exists and where they do not match. We do this comparison for free at the top level, without access to the admin.
What we measure
Current structureHow many categories, how many nesting levels, how many sections are empty. We found a catalog where two of the five root categories did not contain any products.
Directory size by two methodsCMS standard search pagination and site map. In the wholesale catalog we measured, both methods converged on 5,830 products — and it is the discrepancy, when it exists, that shows how many positions the search engine does not see.
Categories without demandSections whose names no one is looking for. They don't hurt, but they don't work either - and the bypass is spent on them in the same way as on working ones.
Uncategorized demandThe most valuable part: groups of requests for which you do not have any page, although the product is in stock.
Competing categoriesSections with almost the same content that share the same group of requests and the same signals.
Category namesAre the sections named as they are searched or as they are named in your accounting system.
Internal search dataWhat people are looking for on the site itself. The cheapest and most honest source of demand, and you already have it - provided that search is enabled and queries are written to analytics.
What you get
A map of the current structure with empty and duplicate partitions marked.
A list of demand areas for which you do not have pages.
Estimate the amount of full work by the number of clusters, not by the number of requests.
Conversation for 40 minutes on the document.
Timeline: 3-5 working days
Why is it free
Because at the top level, the discrepancy between structure and demand is quickly visible, and this is enough to understand whether it makes sense to work full time. If your structure is in order, but sales are falling, we will say so - the document remains with you in any case.
What's next
Next is a full breakdown with clustering and a plan for structural changes. We immediately tell you which changes require changes to the CMS and transitions from old addresses, and which can be done with new pages.
Short form: your contact and site URL
Did not find your case?
Describe how it works on your side — we will tell you whether “Semantics for the directory structure” fits and what it means in your situation. No brief and no call: one question, one answer.
Contract, act and 30-day warranty
Every project gets a written contract: scope, deadlines, amount, acceptance procedure. After delivery — act and invoice, then 30 calendar days of warranty.
Sole proprietor & bank transfer
The contractor is a registered sole proprietor. Payment by invoice with closing documents.
Rights & access — yours
Code, design and materials transfer to you after full payment. Domain, hosting, repository and analytics are registered to you.
Client portal instead of email chains
During the project you get access to a portal: contracts, invoices, acts and project status in one place.
European clients
Among our work — projects for Norway, Bulgaria, Moldova and Spain.
Verifiable numbers
Every case in the portfolio comes with a link to a live site and a technical measurement.
Audit first, then pricing
There is no price list on the site intentionally: the scope of the same work differs multiples between clients.
We say "no" when unsure
If the task isn't ours or the deadline is unrealistic — we tell you upfront.
What affects the price
Why two seemingly identical tasks are priced differently
Width of assortment and number of directionsA catalog with seven sections and a catalog with six hundred categories are different in terms of time-consuming parsing, even with the same number of requests in the file. Volume is defined by width, not frequency.
Depth of nestingThe spare parts catalog we measured holds 3,868 category nodes for 55,321 product addresses. Each level is a separate layer of decisions: what to take up, what to leave inside, where to stop.
Catalog download statusPrice with filled characteristics allows you to match clusters with items semi-automatically. The price, where the characteristics are sewn into the product name in a row, must be disassembled by hand - this is a separate stage.
Number of language versionsThe second language is not a core translation, but a separate demand: people search with other words and brands. In the wholesale catalog from our measurement, each language card contains 6,234 addresses, and the clusters under them are counted separately.
Is there internal search data and Search ConsoleWhen both sources are available, part of the demand is delivered quickly and for free. When they are not there, hypotheses have to be built from the assortment and checked by issuing — longer and with a greater error.
Is TK required for implementation?If your team starts the structure, we write a task with a list of addresses and transitions. When there are many changes and they affect existing addresses, this is a separate document, not a paragraph in the letter.
Cases
Tasks and results in numbers — all metrics measured by us
wholesale supplier of goods from China, B2B catalog of almost 6,000 items, two language versions
Task
Understand how many pages are actually in the directory and whether this number matches what the search engine sees.
Solution
We counted the catalog using two independent methods — the pagination of the regular CMS search and the site map — and separately compiled commercial and non-commercial addresses for each language version.
Result
5,830 products: 58 search pages for 100 items plus 30 for the last one, and the same figure is confirmed by the site map. There are 6,234 addresses in each of the two language cards, for a total of 12,468. The difference between 6,234 and 5,830 is 404 addresses that are not product cards: categories and information pages. That is, approximately for every 14 items in the catalog there is one non-commodity node. The breadcrumbs are passed in the BreadcrumbList markup, so the structure is given to the search engine explicitly, rather than being guessed from the links. Measured on 07/31/2026.
online store of garden equipment and spare parts, catalog of over 50,000 items
Task
Calculate how extensive the category structure of the spare parts catalog is and whether it can withstand such a volume.
Solution
Bypassing the submaps of the site with a separate count of product and category addresses: product submaps — 18 submaps with 3,000 entries plus 1,321 in the nineteenth, category — 3,000 plus 868 in the second.
Result
55,321 product addresses and 3,868 category addresses, for a total of 59,189. A separate category node occurs about every 14 items—the same proportion as in a wholesale catalog ten times smaller. The structure is built for the exact search of a spare part, and it is this that keeps a catalog of this volume navigable. The price of this can be seen separately: the first response of the server is 840.5 ms and the full HTML download is 997.4 ms — expected for 55 thousand products and a point of technical work, not semantic. Measured on 07/31/2026.
What we need from you
We can't start without this — best to prepare in advance
1Downloading the catalog with categories and characteristics — in CSV, XML or by downloading from the accounting system.
2Access to Search Console: part of the answer to your request is already there, and we will not invent it for you.
3Internal search data on the site, if it is configured, or consent to enable the recording of queries in analytics.
4Understanding marginality by direction: priorities should be based on money, not on the number of requests.
5Structure decision maker: Renaming a category is a business decision, not a contractor decision.
If something is missing — let us know, we'll help you gather it or do it as a separate task.
FAQ
Most frequently asked questions — with concrete answers
Why semantics, if the catalog is already made?
Because the catalog is almost always made according to the logic of accounting, and not according to the logic of demand. Sections are named as the vendor names them, and are searched for in other words — sometimes without any common root. The job isn't to put together a query file, but to make your list of pages match the list of what people are asking about. A fair limit: the restructuring of the demand structure itself does not add. If the product is not searched for, the new category under it will remain empty — it will simply cease to be visible to those who are still searching.
How many requests should there be in the kernel?
This is an incorrect unit of measurement. The value is the number of clusters — groups of requests, each of which is served by one page. A thousand requests can be reduced to forty pages, and precisely forty is the number by which the work, term and further filling are calculated. When the volume is called in thousands of requests, the size of the file is called, not the size of the task.
Where do you get the demand from if I have almost no traffic?
From three sources, and none of them need a lot of traffic. Search Console shows what you're already being shown for, even if it's not much. An internal search on the site shows that people are asking with words inside the catalog: the person is already with you, they want something specific and did not find it in the menu. A third source is the structure of the niche: what sections even exist among those who sell the same thing. Our own calculations of the catalogs were made by the regular search of the CMS, so how informative this source is, we do not know from other people's words. We will mention the limitations as well: the internal search is silent where it is turned off or where queries are not written to analytics, and it will not tell anything about a person who still chooses and does not know the name.
What to do with categories that do not have a product?
Decide on each one separately: fill, reduce with the neighboring one or remove with a transition to the nearest living one. An empty section is a dead end for the buyer and a page with no content for the search engine, and the traversal on it is spent the same as on the working one. We bring out such sections in a separate list already at the free verification: this is the fastest part of the work and the cheapest to fix.
How many categories is normal for my catalog?
There is no universal norm, but there is a reference point based on one's own measurements. In the wholesale catalog for 5,830 items of non-commercial addresses, 404 were found, in the spare parts store for 55,321 items — 3,868 category nodes. In both cases, approximately one node per 14 products is obtained, although the catalogs themselves differ in size tenfold. These are two dimensions, not a regularity, and we will not pass them off as the market norm. But here's how the sanity check works: If you have a few hundred items in one category, you're almost certainly lacking nesting—the buyer is scrolling instead of selecting.
Can categories be renamed without consequence?
Without consequences - no. Renaming usually entails a change of address, and the old address may have already collected links and traffic, so transitions from old to new are required. Because of this, we plan structure changes in batches and together with the address match map, rather than one category per week. Separately, we check a technical detail: in the CMS part, the category name and the address segment are resolved, that is, the name can be changed without touching the address at all. Is it so in your system - we find out before planning anything.
Where to start if hundreds of pages are missing?
From the intersection of demand and your margin, not from the volume of requests. A category with high demand and zero margin is a spent budget on content, texts and links. That is why we ask for data on the profitability of directions: without them, the queue is built only by the number of requests, which is not the same as by money. In practice, it looks like a queue of two or three dozen clusters, where the first five are launched in the first month, and the rest are tightened as texts and goods appear.
Will the new structure create thousands of extra pages?
Maybe if you confuse structure with filters. Combinations of characteristics multiply combinatorially, and without canonical addresses and rules in robots, you will end up with thousands of almost identical pages instead of a few strong ones. Therefore, in the project structure, we separate three things separately: what becomes a full-fledged category with its own content, what remains as a filter without a separate address, and what is brought to the landing page under a specific combination. This decision is made before implementation, because analyzing such addresses retrospectively is a separate job with a different budget.
Submit a catalog download - we'll show you where your structure diverges from demand.
The answer is a map of the current structure with empty and duplicated sections, a list of demand directions without pages, and an estimate of the scope of work by the number of clusters.
From measured cases5,830 products: 58 search pages for 100 items plus 30 for the last one, and the same figure is confirmed by the site map
There is no price list on the site on purpose: the same work differs several times over between two clients, and a “from” figure explains nothing in that case. First a free audit — we count your pages, duplicates and speed — then we name the sum and the deadline and fix both in the contract.