What is vocabulary diversity and what does this Vocabulary Diversity Analyzer do?
Vocabulary diversity describes how many distinct normalized words appear relative to the amount of text. This analyzer reports raw type-token ratio, unique-word percentage, hapax terms, repeated-word share and a 50-word Moving-Average Type-Token Ratio when enough text is present.
How to use Vocabulary Diversity Analyzer
Paste a representative passage.
Choose whether common words should be omitted from the top content-word list.
Review total, unique, TTR, MATTR and hapax metrics together.
Read the text-length caution before comparing the result with another passage.
Key features
- Total and unique normalized word counts
- Type-token ratio and unique-word percentage
- Hapax legomena count and share
- MATTR using a documented 50-word window
- Repeated-word concentration and top content words
Input guide
Words are compared case-insensitively after consistent Unicode normalization. Punctuation does not create vocabulary items, while contractions and hyphenated words remain intact. MATTR appears only for samples of at least 50 words; raw TTR should be compared cautiously across lengths.
How to understand your results
TTR divides unique words by total words. Hapax terms occur once. Repeated Word Share counts tokens belonging to words used more than once. MATTR averages TTR across overlapping 50-word windows, reducing—but not eliminating—the effect of text length.
Worked example
cat dog cat bird
There are four tokens, three unique words and two hapax terms. The TTR is 0.75; MATTR is withheld because the sample is shorter than 50 words.
Common use cases
- Compare lexical patterns in similarly sized drafts
- Explore vocabulary in learner or research text
- Review whether a passage relies heavily on a few terms
- Pair repetition evidence with a broader diversity measure
Vocabulary Diversity Analyzer guidance and best practices
- Compare texts with similar purpose, language and length.
- Keep necessary technical terminology even when it lowers diversity.
- Use MATTR and raw counts together rather than treating one number as quality.
Limitations, assumptions and cautions
- Normalization does not perform full linguistic lemmatization.
- Named entities, quotations and spelling variants affect unique-word totals.
- Higher diversity is not automatically clearer, more accurate or better writing.
Privacy and processing
Runs in your browser.
Runs in your browser. Tokens and lexical statistics stay on the device.
Related tools
These live writing tools use the same consistent text foundation for adjacent questions.
Vocabulary Diversity Analyzer FAQs
Clear guidance for choosing and using the catalogue responsibly.
What is vocabulary diversity?
It describes the range of distinct words used in relation to the total amount of text.
What is a type-token ratio?
TTR is the number of unique normalized words divided by the total word count.
Why does TTR depend on text length?
Longer texts naturally repeat common vocabulary, so raw TTR often falls as the sample grows.
What is Moving-Average Type-Token Ratio?
MATTR averages TTR across overlapping fixed-size windows, here 50 words, to reduce sensitivity to total text length.
What are hapax legomena?
They are normalized words that occur exactly once in the analyzed passage.
Does a higher lexical-diversity value always mean better writing?
No. Clear writing may repeat necessary terms, while excessive variation can reduce precision.
Are capitalization differences treated as separate words?
No. Capitalization is normalized for vocabulary comparison.
Can common stop words be excluded?
They can be excluded from the top content-word list; core TTR and MATTR remain based on all normalized words for consistency.
Can I use this for essays, articles or research writing?
Yes, especially when comparing similar passages and interpreting technical vocabulary carefully.
Does my text remain in my browser?
Yes. Lexical analysis runs locally.
