Hyphens and commas do not split words

The implementation trims both ends with trim(), then splits on /\s+/u. Empty or whitespace-only input counts as 0. Spaces, tabs and line breaks separate chunks; hyphens, commas and apostrophes alone do not. JavaScript \s matches the specification’s WhiteSpace and LineTerminator characters.

That makes state-of-the-art 1, one,two 1, and one, two 2. don’t stop also has two chunks. Punctuation-only ... counts as 1, so do not read this as the number of meaningful words. In the code below, words is the tool’s word statistic.

countText("state-of-the-art").words // 1
countText("one,two").words          // 1
countText("one, two").words         // 2
countText("don’t stop").words       // 2
countText("...").words              // 1
one,two counts as 1; one, two as 2; state-of-the-art as 1
An original diagram based on actual countText results. Quotes mark input boundaries and are not input characters. Numbers are the tool’s counts, not linguistic word counts.

ECMAScript: CharacterClassEscape

Text without spaces is not segmented by language

你好世界 counts as 1 in this tool. That does not mean the Chinese sentence has one linguistic word. Unicode UAX #29 explains that word boundaries are not limited to whitespace and punctuation, and some languages need tailoring. Moyoutil does not apply that word-boundary algorithm or dictionary-based analysis.

countText("你好世界").words // 1
countText("one  two").words // 2
countText("one\ntwo").words // 2
countText(" \t\n").words // 0

Unicode UAX #29: Word Boundaries

Paste text into the character and byte counter to update statistics immediately. Check a representative sentence with hyphens or line breaks before pasting your entire draft. If a submission specifies an editor or counting method, use it for the final check. Adding unnecessary spaces just to change the count changes the text itself.

Frequently Asked Questions

Do repeated spaces or line breaks increase the word count?

Consecutive separators form one boundary. Two spaces between one and two still give 2, and a line break also gives 2. Equal word counts can still have different character counts including spaces or UTF-8 byte counts.