A file can include EF BB BF

Unicode documents EF BB BF as the UTF-8 BOM. In UTF-8 it acts as an encoding signature, not a byte-order switch. Our six-byte file combines those three bytes with abc’s 61 62 63. Moyoutil’s file mode reads original bytes with arrayBuffer; text mode encodes the string as UTF-8 with TextEncoder.

Six bytes containing UTF-8 BOM EF BB BF and abc’s 61 62 63, compared with three bytes without the BOM
Original diagram of constructed example bytes. Each box is one byte: six in the first row and three in the second. This is not a file screenshot or actual user data.

Unicode: UTF & BOM FAQ

ignoreBOM: true retains the character

By default, TextDecoder omits an initial BOM from the output string. This input produces abc. Setting ignoreBOM to true bypasses BOM handling, leaving U+FEFF in the string. Re-encoding the default output gives three bytes; its SHA-256 differs from the six-byte original in this example. Do not assume every editor or copying process behaves this way.

const bytes = Uint8Array.of(0xEF, 0xBB, 0xBF, 0x61, 0x62, 0x63);
new TextDecoder().decode(bytes) === "abc" // true
new TextDecoder("utf-8", {ignoreBOM: true})
  .decode(bytes) === "\uFEFFabc" // true
new TextEncoder().encode(new TextDecoder().decode(bytes)).length // 3

WHATWG: TextDecoder

To compare an original file’s checksum, use Moyoutil’s file mode. Pasting the visible abc hashes a different input. If U+FEFF remains in the string, text encoding includes its bytes. Do not remove the original BOM just to match a checksum; check the reference algorithm and file version.

new TextDecoder().decode(Uint8Array.of(0x61, 0xEF, 0xBB, 0xBF))
  === "a\uFEFF" // true

Frequently Asked Questions

Does it also remove U+FEFF in the middle?

This TextDecoder example omits only the initial BOM. An internal U+FEFF remains; ignoreBOM: true also retains the initial character. We executed the code below and empty input to check these boundaries.