Practical guide and verification
Define whether you need full URLs or unique domains
A document can contain many pages from the same host. Full-URL output preserves path, query, and fragment detail; domain-only output answers a different question by collapsing links to their hosts.
Deduplicate after normalizing the intended representation
Two strings can differ only by case, trailing punctuation, or presentation while referring to the same resource. Decide whether exact-string uniqueness or normalized-domain uniqueness is appropriate before removing duplicates.
Filter HTTPS only when the task requires it
An HTTP link can be intentionally present in old documentation, local devices, or test data. An HTTPS-only filter is useful for security review, but deleting HTTP entries can hide evidence if the goal is inventory rather than cleanup.
Extraction is not reachability or safety verification
Finding a syntactically plausible URL does not prove that it resolves, is safe to open, or belongs to the organization implied by its text. Treat the extracted list as data to review, not a trust decision.
Keep output shape compatible with the next tool
Use newline output for manual review, domain-only mode for host counts, and sorted unique output for deterministic comparison. Export or copy only after confirming that the chosen representation preserves the fields needed downstream.