HTML UtilitiesAPI reference for HTML parsing, serializing and formatting utility functionscodeAPI Reference
Categories

HTML Utilities

Functions for reading html into plain nodes and writing it back with no DOM, and for indenting text and markup.

indentLines

Adds consistent indentation to every line of text.

Syntax

indentLines(text, spaces = 2)

Parameters

Name Type Description
text string The text to indent
spaces number Number of spaces to indent each line (default: 2)

Returns

Indented text. Returns empty string for non-string input.

Example

indentHTML

Indents HTML markup with nesting awareness.

Syntax

indentHTML(html, { indent = ' ', startLevel = 0, trimEmptyLines = true } = {})

Parameters

Name Type Description
html string The HTML string to indent
options object Optional configuration

Options

Name Type Default Description
indent string ’ ’ String to use for each level of indentation
startLevel number 0 Initial nesting level to start indentation from
trimEmptyLines boolean true Whether to remove empty lines from the output

Returns

Indented HTML. Returns empty string for non-string input.

Handles void elements (img, br, hr, input), self-closing tags (<component />), comments, and same-line tags (<tag>content</tag>).

Example

parseHTML

Reads html, a fragment or a whole document, into plain nodes with source spans. Everything stays as written: tag and attribute case, quote style, attribute order, entities, whitespace and comments. A void element takes no children, a raw text element (script, style, textarea, title and the rest of the spec set) holds its content as one text child, and a close tag closes the nearest open element of its name. There is no browser recovery beyond that: no implied tbody, no auto-closed p, no invented html or body. Malformed input never throws and never loses bytes, a tag cut off by the end of the string and a close tag with no open element both stay text.

Syntax

parseHTML(html, { closeOnSlash = true } = {})

Parameters

Name Type Description
html string The html to read
options object Optional configuration

Options

Name Type Default Description
closeOnSlash boolean true Close a non-void element at /> the way template, svg and custom element authors mean it. false reads /> the way a browser does, as an open tag whose content follows

Returns

An array of the top-level nodes in order, an empty array for non-string input. Four node shapes appear in a tree, every one carrying start and end source offsets (end exclusive) so html.slice(node.start, node.end) is the node exactly as written:

Node Shape
element { type: 'element', name, attributes, children, selfClosing, start, end, innerStart, innerEnd }, innerStart and innerEnd bracketing the content between the tags
text { type: 'text', value, start, end }, whitespace runs included
comment { type: 'comment', value, start, end }, value being everything between <!-- and -->
doctype { type: 'doctype', value, start, end }, value being everything between <! and >

An attribute is { name, value, quote }: the name as written (@click, :value, viewBox all survive), the value as written with entities undecoded, '' for alt="" and null for a valueless attribute like disabled, and the quote character used, '' when there was none.

Example

stringifyHTML

Writes nodes back to html. Names and values are written as stored with no escaping, attributes separated by single spaces, a self-closing element as <name />, a void element with no close tag, and every other element with a close tag respelled from its open tag. For ordinary markup stringifyHTML(parseHTML(html)) is html byte for byte, the whitespace inside a tag being the one thing normalized. A value set by hand is written verbatim, so quoting it is the caller’s.

Syntax

stringifyHTML(nodes)

Parameters

Name Type Description
nodes object, array A node or a list of nodes from parseHTML, or built by hand

Returns

The html. Returns an empty string for anything that is not a node or a list of nodes.

Example

voidElements

A Set of the elements that never take children or a close tag, per the HTML spec: area base br col embed hr img input link meta param source track wbr. The one copy every scanner in the framework reads.

rawTextElements

A Set of the elements whose content is text and never markup, per the HTML spec: script style textarea title xmp iframe noembed noframes plaintext.

Previous
Functions
Next
Looping