1. Five-second answer: it looks like one rabbit, but it is a three-code-point stack
Meet the creature:
ꪔ̤̮
It looks suspiciously like a tiny rabbit staring back at you. Unicode, however, does not encode this sequence as a rabbit emoji. Its base is U+AA94 TAI VIET LETTER LOW TO, followed by U+0324 COMBINING DIAERESIS BELOW and U+032E COMBINING BREVE BELOW.[1][2]
So the visual result says “small woodland animal,” while the internal representation says “Tai Viet letter plus two combining marks.”
Cute face. Complicated paperwork.
2. Thirty-second story: put it in the “grass,” and it starts living there
The joke began with a decorative kaomoji:
- ̗̀ ꪔ̤̥ꪔ̤̮ꪔ̤̫ ̖́-
Then somebody placed one of the rabbit-looking clusters inside Myanmar text:
မြန်မာမြန်မာမြန်မာ ꪔ̤̮ မြန်မာမြန်မာ
To a reader who does not know the script, the rounded shapes can look like dense vegetation, and the face suddenly resembles a rabbit peeking out of it.
Thai was even more effective:
ภาษาไทยꪔ̤̮ภาษาไทย
The immediate reaction becomes: “Wait. Is there a rabbit in there?”
Naturally, the only responsible next step was to add ten.
ภาษาไทยꪔ̤̮ภาษาꪔ̤̥ไทยภาษาꪔ̤̫ไทยꪔ̤̮กำลังꪔ̤̥เขียนอยู่ꪔ̤̫ภาษาไทยꪔ̤̮ภาษาꪔ̤̥ไทยภาษาꪔ̤̫ไทยꪔ̤̮
At that point, you are no longer reading text. You are conducting a wildlife census.
3. What the “rabbit” actually contains
The three common variants in this joke break down as follows:
| Appearance | Code points |
|---|---|
ꪔ̤̥ |
U+AA94 + U+0324 + U+0325 |
ꪔ̤̮ |
U+AA94 + U+0324 + U+032E |
ꪔ̤̫ |
U+AA94 + U+0324 + U+032B |
U+AA94 is a real consonant in the Tai Viet script.[1] The marks placed under it are combining diacritics with their own Unicode identities.[2]
Unicode did not secretly encode a rabbit here. Human visual perception simply found a face in a legal sequence of characters.
Kaomoji culture saw an opening and took it.
4. Why Thai makes such good camouflage
This is not a claim that Thai script is “weird.” It is a visual joke from the perspective of someone who cannot read it.
Thai writing uses signs that can appear around a base character, including above and below it. Ordinary Thai prose also does not separate every word with spaces the way English does. Unicode’s line-breaking specification therefore classifies Thai among scripts that need language-dependent context analysis to find appropriate break opportunities.[4]
Now insert ꪔ̤̮, which itself places two combining marks below a Tai Viet base.
A Thai reader can immediately see the foreign character sequence. A non-reader, however, may perceive the entire line as a continuous field of curves and marks. The rabbit shape gains accidental camouflage.
ภาษาไทยꪔ̤̮ภาษาไทย
Translation: not linguistic translation, but visual translation — grass, grass, grass… why is something looking at me?
5. Myanmar, Khmer, Malayalam, and Sinhala can join the same visual joke
The effect also works with several visually dense scripts:
Myanmar:
မြန်မာမြန်မာ ꪔ̤̮ မြန်မာမြန်မာ
Khmer:
ខ្មែរខ្មែរអក្សរ ꪔ̤̮ ខ្មែរខ្មែរ
Malayalam:
മലയാളം മലയാളം ꪔ̤̮ മലയാളം
Sinhala:
සිංහල භාෂාව ꪔ̤̮ සිංහල
Thai:
ภาษาไทยกำลังเขียนอยู่ ꪔ̤̮ ภาษาไทย
Any “ranking” here is purely subjective. The scripts are legitimate writing systems, not decorative backgrounds. The game is simply: if you do not read the script and your brain treats the shapes as texture, how well does the rabbit-face cluster blend in?
Thai is an absurdly strong contestant.
6. Then translation mangled the line — but normalization is not automatically guilty
After the ten-rabbit string went through a translation-style workflow, a malformed display with unexpected spaces was observed:
ภาษาไทยꪔ̤̮ภาษาꪔ̤̥ไทยภ าษาꪔ̤̫ไทยꪔ̤̮กำลังꪔ̤̥เข ียนอยู่...
It is tempting to say, “Unicode normalization broke it.” That is too specific without evidence.
Unicode normalization defines ways to decompose, canonically order, and compose equivalent sequences.[5] A translation product, meanwhile, can also perform sentence segmentation, word-boundary detection, tokenization, unknown-character handling, storage conversion, font fallback, layout, and copy/paste transformations.
Thai and Myanmar already require more context-sensitive segmentation than space-delimited English.[4] Mixing one of those scripts with a Tai Viet base plus multiple combining marks creates a useful stress case for the entire text pipeline.
To find the actual culprit, compare the exact code-point sequence after each stage: input, storage, API request, translation result, response parsing, and rendering.
The rabbit is cute. The incident report is not.
7. A “character” on screen is not necessarily one Unicode code point
Unicode text processing distinguishes code points from grapheme clusters, which approximate what users perceive as individual characters.[3]
A visible accented letter may be encoded as one precomposed code point or as a base plus one or more combining marks. Interfaces are expected to treat user-perceived characters sensibly for cursor movement, deletion, selection, and related operations.
That gives our rabbit several identities at once:
- Visually: one rabbit.
- Internally: three code points.
- In a string-length API: the count depends on the programming model.
- For cursor movement: ideally one grapheme-like unit.
- For Backspace: behavior can differ by implementation.
- For search and comparison: normalization policy matters.
- For fonts: fallback and mark positioning can change the appearance.
The wildlife registry is apparently code-point based.
8. The joke accidentally makes a good Unicode test fixture
This tiny creature can test a surprising amount:
- UTF-8 round trips.
- Preservation of combining marks during copy/paste.
- Code-point versus grapheme-aware length handling.
- Cursor and deletion behavior.
- NFC/NFD comparison policy.
- Font fallback and mark placement.
- Line breaking in Thai, Myanmar, or Khmer contexts.
- Translation API round trips.
- Search indexing.
- Cross-browser and mobile rendering.
Do not scatter it randomly through production content. Keep unusual test strings in explicit fixtures with expected code points and expected rendering behavior.
Do not release wild rabbits into your production database.
9. Practical rules for multilingual sites
The boring fixes are the useful ones:
- Standardize on UTF-8.
- Define an explicit normalization policy; NFC is commonly used for web content, while source preservation may require exceptions.[5]
- Distinguish bytes, code points, code units, and grapheme clusters when counting text.
- Set correct
langmetadata and test appropriate fonts and line height. - Test Thai, Myanmar, Khmer, and other scripts that cannot be handled reliably by naive space splitting.[4]
- Do not silently rewrite source-authored text during translation ingestion.
- QA mark placement, wrapping, selection, copy/paste, search, and accessibility — not just mojibake.
- When corruption appears, compare the string at every layer before blaming “Unicode.”
Unicode usually did what the specification said. The bug often appears when software assumes that “one visible character equals one simple unit.”
That is when the rabbit walks out of the forest.
10. Conclusion: the rabbit was human pattern recognition sitting on top of serious text engineering
The story started with a cute Simeji-style kaomoji.
ꪔ̤̮
Put it in Myanmar text and it looked hidden in grass. Put it in Thai and it seemed to belong there. Add ten and you have a herd. Send the herd through a translation pipeline and suddenly you have a Unicode debugging exercise.
Under the joke are real lessons:
- One visible character can contain multiple code points.
- Combining marks attach to bases.
- Some scripts require language-aware segmentation.
- Normalization, segmentation, line breaking, fonts, translation, and rendering are separate layers.
- A strange screenshot does not identify the guilty layer.
And once you have seen this:
ภาษาไทยꪔ̤̮ภาษาไทย
you cannot unsee it.
We came to study Unicode. We left after confirming rabbit habitat.
