TF-IDF Term Analyzer
A TF-IDF term analyzer compares the words in your draft against the pages already ranking for your target query, then surfaces the terms competitors lean on that your copy under-uses or misses entirely. TF-IDF (term frequency × inverse document frequency) rewards words that are common in a single page but rare across the whole set, which is a far better signal of what a page is genuinely about than a raw word count, and it is the same family of math search engines have used for decades to weight content relevance.
This tool runs the full calculation in your browser. Paste your draft and up to five competitor pages, pick unigrams or unigrams plus bigrams, and it tokenizes every document, strips English stopwords, computes real TF-IDF scores, and ranks a content-gap list of terms to add. It is free, requires no sign-up, and nothing you paste is sent anywhere.
| Term | Your TF-IDF | Avg comp. TF-IDF | Docs | Gap |
|---|
🔒 Private: everything runs in your browser. Nothing you paste is uploaded.
How to use the TF-IDF term analyzer
Paste your draft into the first box, drop two or three competitor pages that already rank for your query into the boxes below, then click Analyze to see which terms close the content gap.
Choose competitors that actually rank for your target query
The analysis is only as good as the corpus you feed it, so pull the competitor copy from pages that already rank on page one for the exact query you are targeting, not from a homepage or a loosely related post. Copy the main body text, skip navigation, cookie banners, and footers, and paste each page into its own box. Three to five strong competitors give the inverse-document-frequency weighting enough signal to separate genuine topic terms from noise. A single competitor will still run, but the more relevant pages you add, the more trustworthy the gap list becomes.
Read the gap list before the full table
The orange terms to add panel is the headline output: these are words competitors use that your draft is missing or using at less than half the corpus average. Each pill shows whether the term is missing entirely or merely under-used, and how many competitor pages contain it. Work these into your draft naturally where they fit the topic, never by stuffing. The full table beneath shows your TF-IDF, the average competitor TF-IDF, and a balanced or over-used flag, so you can also spot terms you are leaning on harder than the pages that already rank.
Switch on bigrams to catch multi-word topics
Unigram mode scores single words, which is fast and surfaces the core vocabulary, but real search topics are often two-word phrases. Switch to the unigrams-plus-bigrams tab and the tool also scores adjacent word pairs such as “aged domain” or “backlink profile,” revealing phrase-level gaps a single-word view hides. Bigrams expand the vocabulary count you see in the summary chips, so expect a longer term list. Use unigrams for a quick vocabulary sweep and bigrams when you want to align the exact phrasing competitors use around your primary entity.
TF-IDF term analyzer frequently asked questions
Q1What is TF-IDF and why does it matter for SEO?
TF-IDF stands for term frequency times inverse document frequency. Term frequency counts how often a word appears in one document, and inverse document frequency lowers the weight of words common across every document. The product highlights terms that define a single page rather than the whole topic. For SEO it is a practical way to compare your draft against ranking pages and find the relevant vocabulary your content is missing.
Q2How many competitor pages should I add?
Three to five competitor pages give the cleanest results. Inverse document frequency needs several documents to tell a genuine topic term apart from a word that simply appears everywhere, so a single competitor produces a coarse score. Pick pages that already rank on page one for your exact target query, paste the main body text of each, and the tool weighs your draft against that ranking set to surface the most useful content-gap terms.
Q3Is this TF-IDF tool really free and private?
Yes. The entire calculation runs in your browser with client-side JavaScript. Your draft and competitor text are never uploaded, logged, or stored, and the tool keeps working offline once the page has loaded. There is no sign-up, no usage limit, and no keyword-volume data is fetched from any external service, so it is safe to use on unpublished drafts and confidential client content.
Q4Should I just add every term the gap list shows?
No. Treat the gap list as candidate vocabulary, not a checklist to stuff. Add only the terms that genuinely fit your angle and read naturally in context, and ignore any that belong to a sub-topic you are not covering. TF-IDF measures relevance, not search demand, so it tells you which words ranking pages emphasize, not how many people search them. Pair it with real keyword research before committing to a term.
Q5Why are common words like the and and missing from the results?
The tool removes a built-in English stopword list and any token shorter than two characters before scoring, so function words such as the, and, is, and of never appear in the table. These words carry no topical signal and would otherwise dominate the term frequency counts. It also drops pure numbers. What remains is the content vocabulary that actually distinguishes one page’s subject from another’s, which is what TF-IDF is designed to weigh.
