How to Update Unicode Tables#
LLVM and Clang use a subset of the Unicode Character Database
(UCD) for identifiers, \N{...} named characters, diagnostic column width,
and simple case folding. The corresponding tables are generated from the UCD
into Clang and LLVM Support.
Downloading the Unicode data#
Download the following files:
https://www.unicode.org/Public/UCD/latest/ucdxml/ucd.nounihan.flat.zip
https://www.unicode.org/Public/UCD/latest/ucd/UnicodeData.txt
https://www.unicode.org/Public/UCD/latest/ucd/NameAliases.txt
https://www.unicode.org/Public/UCD/latest/ucd/extracted/DerivedName.txt
Unzip ucd.nounihan.flat.zip to get ucd.nounihan.flat.xml.
Building the generators#
Build the generators from an LLVM build directory that has utilities enabled
(LLVM_BUILD_UTILS, the default) and libxml2 (LLVM_ENABLE_LIBXML2, also the
default):
ninja -C <build> UnicodeCharSetsGenerator UnicodeNameMappingGenerator
Then run them from the llvm-project root as shown below.
Character properties#
UnicodeCharSetsGenerator writes the identifier character sets used by the
lexer (XID_Start, XID_Continue, and the mathematical compatibility notation
profile) and the printable / formatting / combining / East-Asian-width sets
used by LLVM Support.
<build>/bin/UnicodeCharSetsGenerator ucd.nounihan.flat.xml \
clang/lib/Lex/UnicodeCharSetsGenerated.cpp \
llvm/lib/Support/UnicodeCharSetsGenerated.cpp
clang-format -i clang/lib/Lex/UnicodeCharSetsGenerated.cpp \
llvm/lib/Support/UnicodeCharSetsGenerated.cpp
The C99 and C11 identifier tables and the whitespace table in
clang/lib/Lex/UnicodeCharSets.h are maintained by hand.
Character names#
UnicodeNameMappingGenerator writes the name-to-codepoint trie used by
\N{...} and llvm::sys::unicode::nameToCodepointStrict /
nameToCodepointLooseMatching.
<build>/bin/UnicodeNameMappingGenerator UnicodeData.txt NameAliases.txt \
llvm/lib/Support/UnicodeNameToCodepointGenerated.cpp
clang-format -i llvm/lib/Support/UnicodeNameToCodepointGenerated.cpp
Algorithmically derived names (CJK UNIFIED IDEOGRAPH-*, and so on) are not
in those files. Update GeneratedNamesDataTable in
llvm/lib/Support/UnicodeNameToCodepoint.cpp from DerivedName.txt.
Case folding#
llvm/utils/unicode-case-fold.py fetches CaseFolding.txt and writes
llvm/lib/Support/UnicodeCaseFold.cpp:
llvm/utils/unicode-case-fold.py \
https://www.unicode.org/Public/UCD/latest/ucd/CaseFolding.txt \
> llvm/lib/Support/UnicodeCaseFold.cpp
After regenerating the tables, update
llvm/unittests/Support/UnicodeTest.cpp and clang/test/Lexer/unicode.c for
new characters and changed name ranges.