Practical guide and verification
Decide whether case should distinguish tokens
Case-sensitive counting treats Word and word as different tokens, while many text-analysis tasks want them combined. Choose the setting that matches the question and document it when comparing counts across files or revisions.
Punctuation and apostrophes affect tokenization
Hyphens, apostrophes, URLs and Unicode punctuation can change where one word ends and another begins. Inspect representative tokens from the source when the exact unique count matters instead of assuming every counter uses the same tokenization rules.
Unique count and vocabulary richness are different ideas
A longer document usually contains more unique words simply because it contains more words. If the goal is comparing lexical variety, pair the unique count with total word count or a normalized measure rather than ranking documents by raw unique words alone.
Normalization can intentionally merge forms
Lowercasing, trimming punctuation or stemming may combine tokens that were originally distinct. That can be useful for search, deduplication and vocabulary analysis, but it changes the question being answered. Keep the normalization policy explicit.
Verify with a deliberately small sample
Use a short string such as red blue red BLUE and predict the result before running the tool. Under case-insensitive counting there should be two unique words; under case-sensitive counting there should be three. A tiny known sample is the fastest way to verify the chosen settings.