Part of the ScholarTool ecosystemVisit ScholarTool
Writing & Text · Free local tool

Free Vocabulary Diversity Analyzer

Analyze lexical variety with transparent counts and text-length cautions instead of reducing vocabulary to a simplistic writing-quality score.

Live local analysis

Analyze your text

Results update as you type. Your input is never changed automatically.

Runs in your browser

Editable plain text · tested with at least 2,000 words · normal clipboard and undo behavior supported

0 words0 characters0 sentences0 paragraphs

Runs in your browser. Analysis is deterministic and local. Sharing, citation and rating controls never include your writing.

Current analysis

Vocabulary Diversity Analyzer results

Every value below comes from the text currently in the editor.

Add text to begin

Counts and analysis will appear here without a Calculate button.

Use this result

Result utilities

Downloads are created locally from the current analysis.

Real feedback

Was this tool useful?

Public totals appear only when persistent rating storage is available.

Rate this tool
Was this result helpful?

ScholarTool Vocabulary Diversity Analyzer guide artwork with a constellation of varied lexical tiles
ScholarTool writing guideVocabulary Diversity AnalyzerOriginal vertical artwork for this tool and its public feed entry.

What is vocabulary diversity and what does this Vocabulary Diversity Analyzer do?

Vocabulary diversity describes how many distinct normalized words appear relative to the amount of text. This analyzer reports raw type-token ratio, unique-word percentage, hapax terms, repeated-word share and a 50-word Moving-Average Type-Token Ratio when enough text is present.

How to use Vocabulary Diversity Analyzer

  1. Paste a representative passage.

  2. Choose whether common words should be omitted from the top content-word list.

  3. Review total, unique, TTR, MATTR and hapax metrics together.

  4. Read the text-length caution before comparing the result with another passage.

Key features

  • Total and unique normalized word counts
  • Type-token ratio and unique-word percentage
  • Hapax legomena count and share
  • MATTR using a documented 50-word window
  • Repeated-word concentration and top content words

Input guide

Words are compared case-insensitively after consistent Unicode normalization. Punctuation does not create vocabulary items, while contractions and hyphenated words remain intact. MATTR appears only for samples of at least 50 words; raw TTR should be compared cautiously across lengths.

How to understand your results

TTR divides unique words by total words. Hapax terms occur once. Repeated Word Share counts tokens belonging to words used more than once. MATTR averages TTR across overlapping 50-word windows, reducing—but not eliminating—the effect of text length.

Worked example

Example input
cat dog cat bird
How to read it

There are four tokens, three unique words and two hapax terms. The TTR is 0.75; MATTR is withheld because the sample is shorter than 50 words.

Common use cases

  • Compare lexical patterns in similarly sized drafts
  • Explore vocabulary in learner or research text
  • Review whether a passage relies heavily on a few terms
  • Pair repetition evidence with a broader diversity measure

Vocabulary Diversity Analyzer guidance and best practices

  • Compare texts with similar purpose, language and length.
  • Keep necessary technical terminology even when it lowers diversity.
  • Use MATTR and raw counts together rather than treating one number as quality.

Limitations, assumptions and cautions

  • Normalization does not perform full linguistic lemmatization.
  • Named entities, quotations and spelling variants affect unique-word totals.
  • Higher diversity is not automatically clearer, more accurate or better writing.

Privacy and processing

Runs in your browser.

Runs in your browser. Tokens and lexical statistics stay on the device.

Continue analyzing

These live writing tools use the same consistent text foundation for adjacent questions.

Helpful answers

Vocabulary Diversity Analyzer FAQs

Clear guidance for choosing and using the catalogue responsibly.

What is vocabulary diversity?

It describes the range of distinct words used in relation to the total amount of text.

What is a type-token ratio?

TTR is the number of unique normalized words divided by the total word count.

Why does TTR depend on text length?

Longer texts naturally repeat common vocabulary, so raw TTR often falls as the sample grows.

What is Moving-Average Type-Token Ratio?

MATTR averages TTR across overlapping fixed-size windows, here 50 words, to reduce sensitivity to total text length.

What are hapax legomena?

They are normalized words that occur exactly once in the analyzed passage.

Does a higher lexical-diversity value always mean better writing?

No. Clear writing may repeat necessary terms, while excessive variation can reduce precision.

Are capitalization differences treated as separate words?

No. Capitalization is normalized for vocabulary comparison.

Can common stop words be excluded?

They can be excluded from the top content-word list; core TTR and MATTR remain based on all normalized words for consistency.

Can I use this for essays, articles or research writing?

Yes, especially when comparing similar passages and interpreting technical vocabulary carefully.

Does my text remain in my browser?

Yes. Lexical analysis runs locally.

507 focused utilities

Find the right tool

Search by tool, category or the task you want to complete.

Start typing to explore the complete catalogue.