The same visible character can have different code point sequences
Our first example compares é as U+00E9 with U+0065 followed by U+0301. The second sequence contains two code points: e and a combining accent. Fonts and rendering can affect appearance, so we identify the original sequences explicitly. Unicode treats these representations as canonically equivalent.
Unicode UAX #15: Canonical equivalence and normalization
Character counts, bytes and Base64 measured with the actual tool functions
We checked these inputs with the project’s countText, base64 and digest functions on October 5, 2026. Character count here means the number of code points counted by Moyoutil, not necessarily the number of characters a reader perceives. The tools do not automatically convert these inputs to NFC.
| Code points | Character count | UTF-8 bytes | Base64 |
|---|---|---|---|
U+00E9 | 1 | 2 | w6k= |
U+0065 U+0301 | 2 | 3 | ZcyB |
U+AC00 | 1 | 3 | 6rCA |
U+1100 U+1161 | 2 | 6 | 4YSA4YWh |
The Korean examples compare a precomposed Hangul syllable with a sequence of conjoining jamo. These measurements apply to the listed inputs; they are not fixed byte counts for all Korean or accented text. TextEncoder uses UTF-8.
Hashing uses the input bytes, not the text’s appearance
Passing the UTF-8 bytes of the two é representations to Moyoutil’s SHA-256 function produces the values below. They differ in this example. Text mode does not normalize the input, and file mode uses the file’s bytes directly. Do not assume that hashing pasted text gives the checksum of its entire source file.
U+00E9
4a99557e4033c3539de2eb65472017cad5f9557f7a0625a09f1c3f6e2ba69c4c
U+0065 U+0301
bf12767b0f2a56b2190075bae8169f656e3ce8d6357d4aff184bc6c7ea48f9f6Enter each representation separately in the Character and Byte Counter, then calculate SHA-256 using text mode in the SHA Hash Calculator. Retyping the same visible character might not reproduce its original sequence. The escape sequences in the code below let you create the exact inputs.
Reproduce the NFC comparison while preserving the originals
JavaScript normalize supports NFC, NFD, NFKC and NFKD. This example compares the two é sequences after converting both to NFC. NFC applies canonical decomposition followed by canonical composition where possible. Matching normalized values does not mean the original sequences were identical.
const a = "\u00E9";
const b = "e\u0301";
console.log(a === b); // false
console.log(a.normalize("NFC") === b.normalize("NFC")); // true
console.log(new TextEncoder().encode(a).length); // 2
console.log(new TextEncoder().encode(b).length); // 3ECMAScript: String.prototype.normalize
NFKC also folds compatibility distinctions: it changes the circled digit ① to 1 in this example. Applying it indiscriminately can change distinctions that matter in the original data. Normalization is separate from changing case or removing whitespace.
Unicode UAX #15: Normalization forms and compatibility distinctions
A workflow for investigating byte limits and checksum mismatches
- Keep the original and establish whether you are comparing text values or complete files.
- Check UTF-8 byte counts, leading and trailing spaces, and line breaks separately.
- Check whether the receiving system specifies a normalization form. Do not change the original arbitrarily when no rule is given.
- If both sides agree to normalize text for comparison, apply the same form to copies and check their bytes and hashes again.
- For an original-file integrity check, compare the supplied checksum without editing or normalizing the file.
Moyoutil’s Character and Byte Counter, Base64 tool and SHA Hash Calculator are not normalization editors. They do not provide an automatic normalization button. This code is a technical example and does not define the submission rules of an external service.
Frequently Asked Questions
Does converting to NFC always reduce the byte count?
No. The two examples here become smaller, but that is not a general rule for every input. Measure the actual bytes after conversion.
Should identical-looking passwords be normalized automatically?
This article does not define how an authentication system handles passwords. Do not change the original without following that system’s explicit rules.