Semantic Content Strategy: Building Topic Clusters
There are 240 rows on the screen. The columns are familiar: keyword, monthly search volume, difficulty score, current position. Sitting side by side in those rows are "teeth whitening price", "teeth whitening prices", "how much is teeth whitening", "teeth whitening cost", "teeth whitening fee 2026". Five separate rows, five separate volume figures, five separate cells. But they are not five topics. They are one question: what will this procedure cost me?
The table looks like a strategy document because it has numbers in it. It is not one. It is an inventory. An inventory tells you what you have; it does not tell you what to do with it.
A language model already reads those five rows as a single question. Whichever spelling the user types, the model lands in the same space of meaning and builds its answer from there. Writing five separate pages therefore does not give you five times the chance — it gives you five times the dilution. This article is about how you get from a keyword list to a topic cluster.
Twenty phrasings of the same intent are one question
Keyword-level thinking made sense once: the engine matched the word on the page against the word in the query. Today the matching does not happen at word level, it happens at the level of meaning. What the model asks while it evaluates your page is not "how many times does this term appear" but "how completely does this page close this topic".
The practical test is simple: open the list you have and check whether two rows would get the same answer. If the paragraph you would write for "teeth whitening price" and the one you would write for "how much is teeth whitening" are word for word identical, those are one row. If it needs a different answer — "is teeth whitening harmful", say — it is a separate sub-topic.
Once you have run that filter, a 240-row table usually comes down to 15-25 real questions. The number varies from business to business, but the rate of shrinkage is striking. You have lost nothing: the total volume is still sitting where it was, you simply now know how many pages you actually have to write.
The second filter runs on intent. Is the person asking about this topic looking for information, making a comparison, or ready to buy? Those three are not served well by the same page. Showing a price table to someone who is gathering information will lose them; so will telling someone who is ready to buy about the history of the procedure. The intent split gives you the skeleton of the cluster architecture described below.
Cluster architecture: what the hub does, what the spokes do
A topic cluster (pillar and cluster in the English-language sources) is a two-layer structure. At the center, one page that covers the topic as a whole; around it, pages that each go deep on a single sub-question.
The hub page hands over the map of the topic. It defines the terms, draws the main distinctions, summarizes each sub-heading in a few paragraphs and links to that sub-heading's own page. The hub's job is not to explain everything; it is to let readers see where their own question gets answered. Length is a result of that, never a target.
The spoke pages answer one question, at the depth the person asking it needs. A spoke page should not be "everything about this topic" but "the complete answer to this question". For language models this has a concrete counterpart: when a model answers a question it does not take the whole page, it takes the passage that corresponds to that question. If the passage means something on its own, it gets quoted. We covered the writing side of that mechanism in detail in the guide to content optimization for LLMs.
Which page is the conversion page? Usually neither the hub nor the informational spokes. Conversion happens on the spoke that sits lowest in the funnel: the pricing page, the "how do we start" page, the service page. Putting a form on the hub is tempting because it carries the most traffic; but the person who lands there is usually not at the decision stage yet. The hub's job is to route, the conversion page's job is to close. Building this architecture across a whole site is the internal linking part of SEO services.
There is no single answer to how many pages a cluster should hold. The measure is not a count: every page must have a question of its own that no other page answers. If a page cannot meet that condition, it is not a member of the cluster — it is a section of a page you already have.
Where you get the question pool
The quality of a topic cluster depends on how real the sub-questions are. The "people also ask" list on your tool's screen is a starting point, but a weak one on its own — because everybody is looking at the same list. The sources you have and your competitor does not are these.
Sales call records. If you have recordings or notes from the last 30 calls, pull out the questions the customer asked in the first five minutes. The first five minutes matter because the sales pitch has not started yet; the person asks the real question in their head. Write those questions down word for word and do not tidy them up. The sentence "is this stuff going to wreck my mouth" is more valuable data than "side effects of teeth whitening".
Outgoing email. In the email your sales team sends there are paragraphs you "explain all the time" — the same explanation copied to one customer after another. That paragraph is an unwritten page. Scanning the last three months of the sent folder and pulling out the repeated explanations takes half a day and usually yields 5-10 page titles.
Support tickets. The pre-sale question says "what am I buying", the post-sale question says "how do I use what I bought". The second one is missing from most content plans, and yet that is exactly where the content that retains an existing customer and reassures a new one lives.
Search suggestions. Type the topic into the search box and read the autocomplete, then search the question again with modifiers like "why", "how", "when", "instead of", "does it". This surfaces the long tail that the volume tool does not hand you.
Asking the models directly. Asking ChatGPT, Gemini or Perplexity "what questions does someone considering this service ask" shows you how the model breaks the topic apart. That is a direct signal about the model's view of the world — there is nothing wrong with using it, but do not take the list as it comes; cross it against your own call notes.
Once the pool is collected, group the questions by meaning. Questions that get the same answer become one row. The groups you are left with are the spoke list for the cluster. Businesses with an ad account can also pull this pool from their query log; we described the method in our article on turning ad data into a content and GEO plan.
The internal link architecture of a cluster
A cluster stands up on its links. There are three directions and all three are needed.
Hub to spoke. As the hub page summarizes each sub-heading, it links to that spoke page. This shows the boundaries of the cluster to the reader and to the crawler alike.
Spoke to hub. Every spoke page carries a link back to the topic as a whole, usually in the opening section or at the end of the page. Without that link the spoke hangs in mid-air and the reader cannot find the next step.
Between siblings. Spoke pages link to each other only where they are genuinely related. Linking from the pricing page to the process page makes sense; an alphabetical "related posts" list does not. Inflating the number of sibling links artificially blurs what the cluster means.
Link text is a heading in its own right. Text like "click here", "more information" or "details" carries zero information about the destination page. Link text should say what the destination page is about. Nor do you have to reach the same destination with word-for-word identical text every time — natural variation works better and does not break the flow of reading.
A common mistake is sprinkling the links in after the article is finished. The link text then reads as if it had been forced into the sentence. The right way is to link while you are writing the sentence that describes the sub-topic. If the link is part of the sentence, it is in the right place.
Content cannibalization: when a second article adds value
Content cannibalization is the situation where more than one of your pages answers the same question and the engine cannot decide which one to show. The result is that both get weaker: links and authority are split across two pages and neither reaches full strength.
A second article adds value if: it speaks to a different intent (information versus purchase), it is written for a different audience (clinic owner versus patient), or it genuinely goes deep on a sub-heading that was a single paragraph in the first article.
A second article does damage if: it repeats the answer to the same question in different words, it was written as an "updated" version while the old one was left in place, or it was produced purely to fill the publishing calendar.
The way to diagnose it is this: in Search Console, filter by search query and look at which pages are being shown for a single query. If two or three of your pages appear in rotation for the same query, you have cannibalization. This measurement has a limit: Search Console shows classic search, not the source selection inside AI answers. It is still the most concrete data you have for diagnosing cannibalization.
Once you have decided, the merge works like this: pick the stronger one — the one with more links, the one that has been live longer, or simply the better-written one — as the main page. Move the original sections out of the other one and into it; do not copy, add what is genuinely missing. Then put a 301 redirect from the weak page to the strong one and update the internal links so they point at the new destination. Deleting the weak page and leaving a 404 means throwing away the links it had earned.
Coverage depth is not the same as word count
"Long content ranks better" was a correlation, not a cause. Long pages ranked well because they closed the topic more completely. Writers who took the word count as the target inverted that relationship and produced 3,000-word empty pages.
The signs of empty length are familiar: three paragraphs of warm-up before the answer arrives, filler sentences of the "this topic has gained importance in recent years" variety, the same idea said twice in different words, definition paragraphs that have nothing to do with the subject.
For models the cost of this is direct: when a model answers a question it is looking for a quotable passage on the page. An answer buried in filler text does not form a passage that means anything on its own. We went through this and similar writing traps one by one in the mistakes people make writing content for AI.
The right measure is coverage, not words: after reading the page, does the person who asked the question it targets still need a second source? If they do not, the length is sufficient — it might be 700 words, it might be 2,500.
Measuring the cluster: look at the group, not the page
Once the cluster is built, measuring it page by page is misleading. A spoke page's traffic may fall while the hub page's rises; that is not a loss, it is the internal routing doing its job.
To look at it cluster by cluster you need three things. First, the list of URLs that belong to the cluster and their combined impressions and clicks — Search Console lets you build a group with a URL filter. Second, the list of questions the cluster covers and whether or not you are cited as a source on each one; this measurement can be done by hand or with a tool, but either way it needs a repeatable question set. Third, which page in the cluster produces the conversions.
The real question to follow is this: when someone asks about this topic, are you cited as a source? That is a different metric from ranking and it has to be tracked directly. How brand and entity information is held inside the models is a field of work in its own right; we cover that side in the guide to knowledge graph and entity management.
One warning: cluster measurement breeds impatience in the short term. As the spokes go live one at a time, none of them shows a dramatic jump on its own; the effect appears all at once after most of the cluster is published. That is why calling it "not working" after the first two pages is a common and expensive mistake.
What a three-page cluster looks like
A small business has neither the budget nor the patience for a 20-page cluster. The good news: three pages is also a cluster.
Here is how it is built. One hub page: the page that describes your main service as a whole and summarizes the sub-headings. Two spoke pages: the two questions that repeat most often in your question pool. Usually one of them is price or cost and the other is process or a safety concern.
Six links: two from the hub down to the spokes, two from the spokes back up to the hub, and two between the spoke pages in both directions. That is all.
When this trio starts to work — that is, when you start appearing in your own questions — you add a fourth page. The criterion for picking the next page is not volume; it is the most frequent question still sitting unanswered in the pool.
For businesses serving a local area the cluster has a geographic axis as well; we described how to build city pages without stamping a name into a template in the article on city-based service pages. If you want outside help mapping your own topic cluster, you can take a look at our solutions.
The cluster logic works the same way on category and programme pages; the sector-specific version of it is on our e-commerce and education pages.
Do not delete the 240-row table you have — just stop mistaking it for a plan. The plan is the 15 questions you pulled out of it. We described how to run the same question list alongside the social media calendar in the article on repurposing blog content for social media.
Frequently Asked Questions
What is the difference between a topic cluster and a keyword list?
A keyword list is an inventory of the different phrasings typed into the search box, and it holds dozens of rows expressing the same question in different forms. A topic cluster is an architecture that groups those rows by meaning and assigns one page to each group: a hub page in the middle covering the topic as a whole, spoke pages around it that each fully answer a single sub-question, and an internal link network tying them together. The list tells you what you have; the cluster tells you what to write.
How many pages should a topic cluster have?
There is no fixed number; the measure is not the page count but whether every page has a question of its own that no other page answers. If a page cannot meet that condition it is not really an independent page but a section of a page you already have, and publishing it separately weakens both of them. For a small business a three-page cluster of one hub and two spokes is a valid start; the scope is widened as the question pool is worked through.
How do I know whether I have content cannibalization?
The most practical diagnosis is to filter by search query in Search Console and look at which of your pages are shown for a single query; if two or three of your pages appear in rotation for the same query, the engine cannot decide which one to show. The limit of this measurement is that Search Console shows only classic search results and does not cover the source selection inside AI answers. Even so, it is the most concrete data you have for diagnosing cannibalization.
Should I delete the old page when I merge two articles?
No, do not delete it. Pick the stronger page as the main one, move the original sections from the weak page into it, and then put a 301 redirect from the weak page to the strong one. Deleting the page and leaving a 404 means throwing away the external links and the accumulated value that page collected over time. As you finish the merge, also update the links given from inside your own site to the old address so they point at the new destination.
Does long content work better in AI search?
Length on its own is not an advantage. Long pages performed well because they closed the topic more completely; pages padded with filler paragraphs to hit a word count do not produce that effect. When a language model answers a question it does not take the whole page but the passage that corresponds to that question and means something on its own; an answer buried in filler text does not form such a passage. The right measure is this: after reading the page, does the person who asked the targeted question still need a second source?