JavaScript String Length Lies About Emoji

JavaScript string length counts UTF-16 code units, not characters. The four-person family emoji has a length of 11, spreads into seven array entries, and reads as exactly one character to the person who typed it. Intl.Segmenter with granularity: "grapheme" counts and cuts what they see, and it has been available in every current browser since April 2024.

Why does JavaScript string length disagree with emoji?

Three units are in play and only the last one matches a human.

const family = String.fromCodePoint(0x1f469, 0x200d, 0x1f469, 0x200d, 0x1f467, 0x200d, 0x1f466);

family.length;      // 11, UTF-16 code units
[...family].length; // 7, code points

const graphemes = new Intl.Segmenter("en", { granularity: "grapheme" });
[...graphemes.segment(family)].length; // 1

That emoji is four people joined by three zero width joiners. Four of its seven code points sit above the Basic Multilingual Plane and take two code units each, which is where 11 comes from.

Spreading the string gets you to code points. Better, still wrong. The GB flag is two regional indicator symbols and spreads into two entries. A thumbs up with a skin tone modifier spreads into two. Devanagari does not improve at all: the Hindi greeting written String.fromCodePoint(0x928, 0x92e, 0x938, 0x94d, 0x924, 0x947) is six code points, six code units, and three characters as far as a Hindi speaker is concerned, because the last three code points form one cluster.

Unicode calls that unit an extended grapheme cluster. UAX #29, the annex that defines the boundaries, is direct about which unit a counter wants: “In those relatively rare circumstances where programmers need to supply end users with user-perceived character counts, the counts should correspond to the number of segments delimited by grapheme cluster boundaries.”

Cutting matters more than counting. family.slice(0, 4) ends halfway through a surrogate pair, and family.slice(0, 4).isWellFormed() returns false. What ships to the page is a replacement box.

What does Intl.Segmenter actually return?

A Segments object, not an array and not an iterator.

const graphemes = new Intl.Segmenter("en-GB", { granularity: "grapheme" });
const segments = graphemes.segment("Gift set for mum");

for (const { segment, index } of segments) {
  // segment: "G", index: 0
  // segment: "i", index: 1
}

segments.containing(11); // { segment: "r", index: 11, input: "Gift set for mum" }

Each iteration hands you a segment data object with segment, index and input. index is a code unit offset into the original string, so it drops straight into slice. containing(index) takes a code unit offset and gives you back the whole cluster sitting at it, which is how you snap an arbitrary cursor position onto a boundary. Out of range, it returns undefined rather than throwing.

The Segments object is re-iterable. Each for...of calls [Symbol.iterator]() and gets a fresh segment iterator over the same string, so unlike the lazy iterator helpers we wrote about, a second pass is not empty.

Granularity is the only option that matters, and ECMA-402 allows three values: "grapheme", "word" and "sentence". The default is "grapheme", so new Intl.Segmenter("en-GB") is already the character segmenter. Ask for anything else, including the "line" you were hoping for, and V8 throws RangeError: Value line out of range for Intl.Segmenter options property granularity.

How do you truncate without cutting a character in half?

Count clusters until you hit the limit, then slice at that cluster’s index.

const graphemes = new Intl.Segmenter("en-GB", { granularity: "grapheme" });

export function truncate(input, limit, suffix = "...") {
  let count = 0;
  for (const { index } of graphemes.segment(input)) {
    if (count === limit) return input.slice(0, index) + suffix;
    count += 1;
  }
  return input;
}

The loop exits at the limit. It never builds an array and never walks the tail of a long string, which is the whole reason for iterating rather than reaching for [...segment(input)].slice(0, limit).join(""). Strings shorter than the limit fall out of the bottom untouched, with no suffix.

Run it over a product title with that family emoji in the middle and a 12 character limit and you get the emoji intact plus the next letter. The same string cut with slice(0, 12) is 12 code units, which lands inside the emoji and renders as a woman plus an invisible joiner.

Counting is the same walk. [...graphemes.segment(title)].length gives 18 for a title whose length is 28. If that number drives a live counter under a textarea, work it out in the input handler and keep it in state. A render that runs on every keystroke should read the number, not recompute it.

When do you want words or sentences instead?

Word granularity adds an isWordLike boolean to each segment, and the spec sets it only for that granularity. Spaces and punctuation come back as segments too, flagged false, which is what makes a word count possible.

const words = new Intl.Segmenter("en-GB", { granularity: "word" });
const text = "It's state-of-the-art, honestly.";

[...words.segment(text)].filter((s) => s.isWordLike).map((s) => s.segment);
// ["It's", "state", "of", "the", "art", "honestly"]

Thirteen segments, six of them word-like. Note what happened to the hyphens: UAX #29 breaks at them, so state-of-the-art counts as four. The apostrophe in It's survives. If your definition of a word disagrees, this is where you stop and write your own rule instead.

The case that earns its keep is text with no spaces. A Japanese sentence of fifteen code units returns one item from split(" ") and ten segments from a ja word segmenter, eight of them word-like. Thai, Khmer, Lao and Chinese have the same problem, and no regex you write is going to fix it, because the boundaries come from a dictionary inside the engine’s ICU data.

Sentence granularity is the weakest of the three. "Dr. Smith went home. He slept." segments into three sentences in English, because the full stop after Dr looks like a terminator. Treat it as a first pass, not as a parser.

What Intl.Segmenter will not do

It will not break lines. Line breaking is a separate Unicode annex, UAX #14, and there is no "line" granularity to ask for. In a browser the layout engine already does it, and CSS gives you hyphens and text-wrap to tune where it happens.

It will not give you stable results across every runtime. Word-like classification is explicitly implementation-dependent in the spec, and the Unicode version behind the boundaries is whatever ICU the engine was built with. In Node, process.versions.unicode tells you which. Two runtimes on different ICU versions can disagree about a brand new emoji sequence, so a grapheme count is not something to hash or store as a checksum.

It will not type check on an old lib. The declarations live in lib.es2022.intl.d.ts, so a tsconfig.json still targeting "lib": ["es2021", "dom"] fails with error TS2339: Property 'Segmenter' does not exist on type 'typeof Intl' even though the runtime is fine.

Support itself is settled. Chrome shipped it in 87 in November 2020 and Safari in 14.1 in April 2021, but Baseline dates the feature to 16 April 2024, when Firefox 125 landed and the last gap closed. Node has had it since 16.0.0. If your support matrix reaches back to Firefox 124, keep whichever of grapheme-splitter or graphemer you already have. Past that line the engine holds the Unicode data itself, and the dependency is doing work the runtime does.

Where to put it, and where not to

Put it anywhere a number or a cut reaches a user. Character counters, product titles truncated to fit a card, meta descriptions trimmed to 155, the preview line in a message list. In all of those, a count of 11 for something that looks like one character is a visible bug, and a bad slice ships a replacement box to production.

Leave it out of validation you control end to end. A slug, an internal ID, a hex token: length is the right check there, and it does not walk the string. The same goes for column limits, where the database counts its own units and the only honest fix is to check against the same unit the database uses. Adopting it is a much smaller decision than adopting Temporal, because there is no dependency, no polyfill and nothing to migrate: one hoisted segmenter and a helper that fits on a screen.

We build ecommerce front ends where this shows up daily, in product titles written by merchandisers who like emoji and in storefronts localised for scripts that do not use spaces. If you want a hand with the multilingual end of a storefront, that is part of our ecommerce development work.

Need this built properly?

Whoooop Ltd has spent 15+ years building and maintaining web applications in TypeScript, React, Node.js and serverless — the same ground this post covers.

Get in touch