html — HyperText Markup Language support¶
Source code: Lib/html/__init__.py
This module defines utilities to manipulate HTML.
- html.escape(s, quote=True)¶
Convert the characters
&,<and>in string s to HTML-safe sequences. Use this if you need to display text that might contain such characters in HTML. If the optional flag quote is true (the default), the characters (") and (') are also translated; this helps for inclusion in an HTML attribute value delimited by quotes, as in<a href="...">. If quote is set to false, the characters (") and (') are not translated.Added in version 3.2.
- html.unescape(s)¶
Convert all named and numeric character references (e.g.
>,>,>) in the string s to the corresponding Unicode characters. This function uses the rules defined by the HTML 5 standard for both valid and invalid character references, and thelist of HTML 5 named character references.Added in version 3.4.
- html.htmlcharrefreplace_errors(exception)¶
Implements the
htmlcharrefreplaceerror handling (for encoding only): the unencodable character is replaced by the corresponding HTML named character reference fromhtml.entities.codepoint2name, or by a numeric character reference if there is no name for it.This error handler is not registered by default, you should register it with
codecs.register_error():>>> import codecs, html >>> codecs.register_error('htmlcharrefreplace', ... html.htmlcharrefreplace_errors) >>> '∀ x∈ℜ'.encode('ascii', 'htmlcharrefreplace') b'∀ x∈ℜ'
Added in version 3.16.0a0 (unreleased).
Submodules in the html package are:
html.parser– HTML/XHTML parser with lenient parsing modehtml.entities– HTML entity definitions