Count the same string four ways

Moyoutil calculates its count including spaces with [...s].length. This counts code points, while JavaScript s.length counts UTF-16 code units. The family sequence contains four person code points and three U+200D joiners, totaling seven. Its UTF-16 length is 11, and TextEncoder produces 25 UTF-8 bytes.

const s = "👨‍👩‍👧‍👦";
[...s].length // 7
s.length // 11
new TextEncoder().encode(s).length // 25
Comparison of one grapheme, seven code points, eleven UTF-16 units, and twenty-five UTF-8 bytes for a family emoji sequence
Diagram of directly executed calculations. Numbers use different counting units; box areas do not represent numerical ratios. It shows counting methods rather than an emoji rendering.

ECMAScript: String iterator · WHATWG: TextEncoder

Text boundaries and input limits

Unicode UAX #29 describes extended grapheme cluster boundaries and rules that avoid breaks within emoji ZWJ sequences. Running this example through Intl.Segmenter with grapheme granularity produces one segment. A grapheme boundary is not a guarantee of actual rendering or the character-count policy of every submission form.

const segmenter = new Intl.Segmenter("en", {
  granularity: "grapheme"
});
[...segmenter.segment(s)].length // 1

Unicode UAX #29: Grapheme Cluster Boundaries

Moyoutil does not provide separate grapheme or UTF-16 counts. Paste this sequence into the character and byte counter to check 7 including spaces, 7 excluding spaces, and 25 UTF-8 bytes. Check whether a limit means characters, bytes, or UTF-16 units, then verify in the destination form. Removing joiners changes the original sequence.

Frequently Asked Questions

Does every emoji count as seven characters?

No. This result applies to the specific family sequence shown here. Different code points produce different counts. An empty string has zero code points, UTF-16 units, UTF-8 bytes, and graphemes.