GlyphSift
Style TextClean TextInspect UnicodeGuidesCompatibilityAll Tools
Open studio ↘
Home/Guides/How to Shorten Text to a Byte Limit Without Breaking Emoji

Unicode guide

How to Shorten Text to a Byte Limit Without Breaking Emoji

Many fields cap length, but they do not all count the same way. Some count characters a reader sees, some count Unicode code points, and some count raw UTF-8 bytes. Trimming safely means measuring in the right unit and never cutting a character in half.

Written & reviewed by Jack Shi · Last reviewed 2026-09-20

Make it now

Generate it in three steps

  1. 1

    Open the Text Length Limiter and paste the text you need to trim.

  2. 2

    Set the maximum length and choose the unit — graphemes, code points, or UTF-8 bytes — to match the destination's rule.

  3. 3

    Turn the ellipsis on if you want a … marker counted inside the budget, then copy the trimmed result.

Open the Text Length Limiter →
01

Three ways to measure length

A grapheme is what a reader perceives as one character, such as é or a family emoji, even when it is built from several code points. A code point is a single Unicode scalar value. A UTF-8 byte is one unit of the encoded form, where a character can take one to four bytes.

A plain ASCII letter is one grapheme, one code point, and one byte, so the three agree. Accented letters, symbols, and emoji make them diverge, which is why the unit you trim by matters.

Inputcafé
GlyphSift output4 graphemes · 4 code points · 5 UTF-8 bytes

The é is one grapheme but two UTF-8 bytes, so a byte limit counts it as two.

02

Keep whole graphemes

A family emoji like 👨‍👩‍👧‍👦 is a single grapheme made of several code points joined by zero-width joiners. Cutting by raw code points or bytes in the middle of that sequence produces broken or unexpected glyphs.

Trimming should therefore stop on a grapheme boundary: add whole characters until the next one would exceed the limit, then stop. That keeps emoji and combined characters intact instead of splitting them.

InputHi 👨‍👩‍👧‍👦 there (limit 4 graphemes)
GlyphSift outputHi 👨‍👩‍👧‍👦

The family emoji counts as one grapheme and is kept whole.

03

Budget the ellipsis

If you add an ellipsis to show text was cut, the ellipsis also takes space. To stay within the limit, subtract its length from the budget first, then fill the rest with content. A single-character ellipsis (…) is one grapheme but three UTF-8 bytes.

Trimming only shortens and reports; it is not a summary. It should tell you the used and total length so you can confirm the result fits the destination's own counting rule.

Inputabcdefg (limit 5 graphemes, ellipsis on)
GlyphSift outputabcd…

Four characters plus the ellipsis fill the 5-grapheme budget.

Try the behavior

Related tools

cleanupText Length Limiter

Trim text to a maximum length in graphemes, code points, or UTF-8 bytes, keeping whole characters and optionally adding an ellipsis.

Open tool →
inspectUTF-8 Byte Counter

Compare UTF-8 bytes, code points, graphemes, and UTF-16 units without uploading the text.

Open tool →
inspectGrapheme Counter

Count user-perceived characters (graphemes) and compare them to code points and UTF-16 units.

Open tool →

FAQ

Frequently asked questions

Why does my text fit one field but not another?
Fields count differently. One may count graphemes a reader sees, another code points, another raw UTF-8 bytes. An accented letter or emoji can be one grapheme but several bytes, so the same text measures differently.
How do I trim text without breaking an emoji?
Trim on a grapheme boundary rather than by raw bytes or code points. The Text Length Limiter adds whole graphemes until the next one would exceed the limit, so emoji and combined characters stay intact.
Does the ellipsis count toward the limit?
It should. Subtract the ellipsis length from the budget first so the final result, including the …, still fits. A single … is one grapheme but three UTF-8 bytes.

Primary standards

References

  • The Unicode Standard, Version 17.0
  • UAX #29: Unicode Text Segmentation (grapheme clusters)
  • RFC 3629: UTF-8, a transformation format of ISO 10646
GlyphSiftStyle it. Clean it. Check it. Paste with confidence.
Style mapGuidesAboutAuthorContactPrivacyTermsEditorial

Independent copy-paste text tools · hello@glyphsift.com · Platform claims require evidence

Static Unicode maps target Unicode 17.0. Segmentation and normalization use the runtime's Unicode data.