Alex Wright Media

How to Cluster Keywords for Topical Authority That Ranks

Most keyword lists are symptoms, not strategies. Here's the method for clustering by SERP overlap—not lexical similarity—to build real topical authority.

How to Cluster Keywords for Topical Authority That Ranks

How to cluster keywords for topical authority (without drowning in 4,000-row spreadsheets)

Open a keyword tool on any mid-sized niche and you'll get back somewhere between 800 and 5,000 rows. That is not a keyword list. That is a symptom. And the reason most people never build real topical authority isn't laziness or budget—it's that they never move past that spreadsheet into actual clusters.

Here's the thing: Google doesn't rank pages one at a time anymore. It ranks topics. When you type "how to cluster keywords for topical authority" into a search bar, you're already halfway there. The hard part is the operational middle—deciding what belongs together, what doesn't, and why.

I've been building content systems like this since roughly 2023. I've made every mistake on this list and a few that aren't. What follows is the method I actually use, plus the parts nobody writes about because they don't fit neatly into a diagram.

Key takeaways

  • Topical authority comes from coverage + interconnection, not from publishing more pages.
  • Group keywords by SERP overlap, not by lexical similarity. Two phrases can look identical and belong to different pages.
  • A cluster of 5-12 pages is usually the sweet spot for a mid-sized niche. Bigger isn't better.
  • Cannibalization isn't a tool feature—it's a clustering decision you make before you publish.
  • Measure clusters by average position per cluster, not per page.

Why keyword clustering and topical authority are the same thing

Most guides treat these as two steps: first you cluster, then you build authority. I'd argue they're one process viewed from different angles. Authority is just what a cluster looks like from Google's side of the fence.

Think about what a cluster actually is. It's a set of queries that share intent and can be satisfied by a set of pages that reference each other. If those pages exist, are internally linked, and answer the full arc of the intent—Google has a reason to treat you as the source for that topic.

The pillar-and-cluster model is fine, but incomplete

The standard model—one pillar page, a bundle of supporting pages, internal links pointing up—is described the same way in almost every guide. It works. I'm not going to pretend otherwise.

What's missing is the part where you decide which keywords go on which page. That decision is where 80% of the outcome lives. Get it wrong and no amount of internal linking saves you.

What Google actually looks for

Two things, in this order:

  • Depth — can you answer the full intent behind a query, including the follow-up questions a user will have next?
  • Breadth — do you cover the adjacent topics someone researching this would reasonably expect to find under the same roof?

Neither one requires a 200-page content library. Both require you to know what "the same roof" means. That's clustering.

How to actually group keywords (step by step)

Forget embeddings and clustering APIs for a second. Start with something primitive: SERP overlap. It's the most reliable signal you have access to without enterprise tooling.

Step 1: start with SERP overlap, not lexical similarity

Two keywords belong to the same cluster when the same pages rank for both. That's it. That's the test.

Take "content marketing strategy" and "content marketing plan." They look like twins. On some queries they rank identically. On others, the first pulls in B2B SaaS blogs and the second pulls in freelance writing templates. Different intent, different cluster.

Pull the top 10 results for each phrase in your list. Compare. If 6 or more URLs overlap between two keywords, they belong together. If 2 or fewer overlap, they don't—no matter how similar the words look.

Step 2: set a similarity threshold and stick to it

Rough threshold I've used successfully: 60% URL overlap = same cluster. Below 40% = separate clusters. Between 40 and 60% = review manually, because that grey zone is where judgment matters.

Automated clustering tools (the ones that do cosine similarity on embeddings) tend to cluster too aggressively. I tried one on a project covering roughly 1,200 keywords and got back 47 clusters. When I reviewed them by hand, about a third should have been merged and another third split. The tool was fast; it just wasn't right.

Step 3: decide cluster size before you build

Here's a rule I've come to trust: one cluster should support 5 to 12 pages. Fewer than 5 and it's not really a topic—it's a page. More than 12 and you're probably looking at two clusters wearing a trench coat.

A quick decision table

  • 1-2 pages: not a cluster, just a topic page.
  • 3-4 pages: thin cluster, often merges with a neighbor.
  • 5-12 pages: the working range.
  • 13+ pages: almost always splits into two clusters once you check SERPs again.

The reason this matters: cluster size shapes your internal linking. Too small and you genuinely don't have enough surface area for meaningful cross-links. Too big and every page competes with every other page for the same query.

The cannibalization nobody plans for

Cannibalization gets mentioned as a feature in keyword tools. It's not a feature. It's a consequence of a clustering decision you've already made—usually a bad one.

The cannibalization nobody plans for

Concrete scenario. I worked on a project where three pages ranked positions 4, 7, and 9 for the same phrase. The team had written them at different times without a cluster plan. Each page was 60-70% overlapping in content.

The fix wasn't to delete two. It was to redraw the clusters: one page got the primary phrase, one got a long-tail variant that had been misassigned, and the third got absorbed as a section on the first page. Positions after four weeks: 2, 14 (a different phrase), and gone (merged). Net gain.

Do that audit before publishing, not after. Ask, for every new page: which existing page in this cluster is this going to compete with? If the answer is "none," good. If the answer is "actually, hmm," reconsider.

Measurement: how do you know clustering worked?

Track these per cluster, not per page:

  • Average position across all pages in the cluster — the leading indicator of authority building.
  • Number of unique queries bringing impressions — if the cluster is working, this number grows even when individual pages plateau.
  • Internal click flow — users who land on a supporting page and click through to the pillar within the same session.
  • New pages ranking faster — a classic authority signal. If your seventh page in a cluster starts ranking within a week while your first took three months, the cluster is functioning.

That last one is the most honest test. On my own project, the first page in a cluster took around 11 weeks to reach page one. The eighth page in the same cluster hit page one in under three weeks. The content wasn't better. The cluster was.

Example: a real cluster, broken down

Concrete, since abstractions are useless here. Niche: project management for small agencies.

Page typeTarget intentInternal links out
Pillar page"project management for agencies"8 supporting pages
Comparison page"best PM tools for creative teams"Pillar + 2 tool reviews
Process guide"how to structure a project brief"Pillar + 1 template page
Template page"project brief template free"Process guide + pillar
Case study"agency PM case study"Pillar + comparison page
Pricing guide"how much to charge for project management"Pillar only
FAQ page"project management questions agencies"Pillar + process guide
Glossary"PM terms for agencies"All pages

Eight pages. One pillar. Clear intent per page. Zero overlap in primary keywords. That's a cluster.

Notice the mix: comparison, process, template, case study, pricing, FAQ, glossary. Not seven blog posts with the same structure. The variety is what makes the cluster feel like coverage rather than content.

What about very small niches?

If your niche only supports 3-4 pages total, you're not building clusters—you're building a topic page and supporting pages. That's fine. Authority happens at the topic level too. Just don't force a cluster where the market doesn't exist.

Common questions about clustering

Do I need paid tools to cluster keywords?

No, but you need patience. Free approaches: pull SERPs manually for 20-30 keywords, build a small overlap matrix in a spreadsheet, apply the 60% rule. It takes a few hours. Paid clustering tools save that time but give you output that still needs review. I've done both—paid saves maybe 70% of the mechanical work and none of the judgment.

Should I cluster by search volume tiers?

No. Volume belongs to your prioritization step, not your clustering step. Clustering is about intent and SERP behavior. If you filter by volume first, you'll break clusters apart that belong together and lose the very interconnectivity that produces authority.

What I would do differently

Looking back at the first clusters I ever built, the mistake wasn't the tooling or the threshold. It was clustering by how keywords looked rather than how they ranked. I remember building a cluster around "email marketing automation" that grouped 14 phrases—almost all of which pulled wildly different SERPs. Produced a mess, ranked nothing for months.

If you take one thing from all this: run the SERP overlap test before you write a single word. It's the cheapest, fastest filter you have. And it catches the errors that no amount of internal linking later can fix.

Which leaves an uncomfortable question worth sitting with. If clustering really is the work that produces topical authority, why is almost every guide still teaching the pillar-and-cluster diagram as though the diagram were the method? The method is the decision about what belongs together. Everything after that is execution.

Marissa Whitaker

Marissa Whitaker is a technical SEO specialist who helps organizations strengthen their search performance through rigorous audits, thoughtful site architecture optimization, and practical improvements to Core Web Vitals. Known for translating complex technical findings into clear, actionable guidance, she works closely with development and content teams to build sites that are fast, crawlable, and easy to navigate. Her approach combines analytical precision with a genuine appreciation for the people who use the sites she helps improve.

See all articles →

Related articles