When str.lower() is a security vulnerability in Python Seth Larson @ 2026-08-18
Some internet standards only support ASCII characters, but the world uses much more than the Latin alphabet. Thus, a mapping from Unicode to ASCII for use in domain names is required.
NamePrep was part of that solution, defined in RFC 3491 as a profile of StringPrep, and is crucially a component of Internationalizing Domain Names in Applications (IDNA), also known as “IDNA 2003”. The StringPrep algorithm is defined in RFC 3454. IDNA 2003 has been obsoleted by IDNA 2008 defined in RFC 5890, 5891, 5892, and 5893.
Python supports IDNA 2003 through the idna codec ( str.encode('idna') ) and IDNA 2008 is supported by the idna package on the Python package Index. Python's implementation of StringPrep is implemented in the stringprep module in the standard library. In general, you should be using the idna package (IDNA 2008) and not .encode("idna") (IDNA 2003), but sometimes you do need the older behavior.
StringPrep defines the “case folding” step (case folding is approximately “how to lowercase/uppercase a codepoint”) in Section 3.2, enabling case-insensitive comparisons of strings, by mapping all characters through mapping tables B.2 and B.3. B.2 is effectively str.lower() , lowercasing all characters according to Unicode rules and B.3 contains the exceptions. The Python code implementing this (and assuming B.3 table is captured correctly) is the following code below:
def map_table_b3 ( code ): r = b3_exceptions . get ( ord ( code )) if r is not None : return r return code . lower ()
And that might seem fine... and the title probably gave it away already. The str.lower() call in this function is a vulnerability!
Why? Because str uses whatever Unicode data that the particular Python interpreter is shipped with, you can figure out what Unicode version your Python interpreter uses by accessing unicodedata.unidata_version :
>>> import unicodedata >>> unicodedata . unidata_version '17.0.0'
There's also a database of Unicode 3.2.0 data available on every version of Python ( unicodedata.ucd_3_2_0 ) specifically for the StringPrep and IDNA algorithms:
... continue reading