Comment 9 for bug 1958539

Revision history for this message
frenzy (frenzy-madness) wrote :

I've prepared this message and I'll start posting it to projects we've found affected in the previous stages.

Subject: Consider switching from lxml's clean_html for enhanced security (and possibly performance)

I'd like to bring to your attention that we are discussing a possibility to remove lxml's clean_html functionality from lxml library. Over the past years, there have been several concerning security vulnerabilities discovered within the lxml library's clean_html functionality – [CVE-2021-43818](https://cve.mitre.org/cgi-bin/cvename.cgi?name=CVE-2021-43818), [CVE-2021-28957](https://cve.mitre.org/cgi-bin/cvename.cgi?name=CVE-2021-28957), [CVE-2020-27783](https://cve.mitre.org/cgi-bin/cvename.cgi?name=CVE-2020-27783), [CVE-2018-19787](https://cve.mitre.org/cgi-bin/cvename.cgi?name=CVE-2018-19787) and [CVE-2014-3146](https://cve.mitre.org/cgi-bin/cvename.cgi?name=CVE-2014-3146).

The main problem is in the design. Because the lxml's clean_html functionality is based on a blocklist, it's hard to keep it up to date with all new possibilities in HTML and JS.

Two viable alternatives worth considering are `bleach` and `nh3`. Here's why:

bleach:
* Bleach is a widely adopted Python library specifically designed for sanitizing and cleaning HTML input.
* It has a strong track record in terms of security – it's allowed-list-based.
* It was deprecated in January but it will still receive security updates, support for new Pythons and bugfixes, see [upstream issue](https://github.com/mozilla/bleach/issues/698#issue-1553425406).

nh3:
* nh3 is Python binding for the ammonia library. Ammonia is written in Rust and it's also allowed-list-based.
* Thanks to the Rust backend, nh3 is also significantly faster than bleach.
* Rust backend is nothing to be afraid of. nh3 uses the latest PyO3 compatible with Python 3.12 and provides wheels built on top of compatible ABI for different architectures and platforms.

We'll probably move the cleaning part of the lxml to a distinct project first so it will still be possible to use it but better is to find a suitable alternative sooner rather than later.

Let me know if we can help you with this transition anyhow and have a nice day.

---

Do you think there is anything to add or adjust?

Lumír