DOM-aware HTML to text: what is preserved
A useful HTML-to-text conversion is not the same operation as deleting everything between angle brackets. Paragraphs, headings, list nesting, links, tables, line breaks, blockquotes and preformatted code all carry structure that can disappear when markup is stripped blindly.
Readable text versus Markdown-ish text
The Readable preset keeps a conventional plain-text layout. Markdown-ish mode uses familiar heading and link notation but remains an extraction format rather than a standards-complete HTML-to-Markdown converter. Choose the output based on where the text goes next, not because one representation is universally more correct.
Tables and links
Tables can become tab-separated rows, CSV rows or pipe-separated lines. Links can keep only their labels, append the href, use bracket-and-parenthesis notation, or output the URL. Supply a base URL when relative links such as /docs/start must become absolute.
Why source resources are neutralized first
HTML parsed as a detached document does not run normal script elements, but parsing can still involve resource-bearing elements in some browser behaviors. The Studio renames load-capable attributes on images, frames, media, objects and stylesheet links before DOM parsing, removes script/style blocks, and never inserts the parsed source tree into the live WebToolArc document.
Batch use
The batch workflow applies the same extraction policy to up to 50 local files and can package text outputs plus a manifest into one ZIP. This is useful for exports, archived pages and CMS migrations where the conversion policy must stay consistent across many documents.