Skip to content
Rohit Behera
← All work
Kind
Own open source · MIT
Year
2026
Role
Author
Stack
PythonUnicode UAXpytestGitHub Actions

splitlint

A CLI and library that audits any RAG text splitter for the Unicode corruption a round-trip test cannot see.

The hard partThe broken chunk still joins back into the original string. Nothing throws; retrieval just gets worse for one language.

  • 766 of 766 official Unicode grapheme conformance cases pass
  • LangChain's RecursiveCharacterTextSplitter fails 10 of its 12 cases
  • Grew out of an upstream fix in langchain4j

The problem

Every retrieval pipeline chunks text before embedding it, and almost every chunker cuts every N "characters". The bug is in what counts as a character.

Count UTF-16 code units, as Java's substring and JavaScript's slice do, and a cut can land between the two halves of a surrogate pair. That chunk can no longer be encoded as UTF-8 at all. Count code points, as Python does, and the cut lands inside a family emoji, a Devanagari conjunct or a Hangul syllable instead. Either way the chunks still join back into the original text, so the obvious test passes. The damaged chunk embeds as something else, and you find out when retrieval quality drops for one script.

I found the first version of this in langchain4j: its splitter worked in UTF-16 code units, and the fix (#6308, in review) splits by code point. splitlint is that investigation turned into a tool that checks any splitter in seconds.

Try it

The same three ways of counting, on text a reader would call ordinary. Change the chunk size and watch where each one cuts.

Splitter · same text, three ways to count

5

This demo runs in your browser. With JavaScript off, the case study below describes the same three failure modes.

Every row joins back into the original text, so a round-trip test passes all three. The broken chunks embed as something else. Grapheme boundaries here come from your browser’s Intl.Segmenter; splitlint ships its own UAX #29 implementation, checked against all 766 official test cases.

How it finds the bug

A family emoji in the middle of a chunk proves nothing. Each of the twelve test cases pads the text with separator-free filler until the hazard straddles the splitter's cut offset exactly. That also pushes recursive splitters past their paragraph, line and space separators, down to the character-level fallback where the boundary bugs live.

Six invariants, graded by what the failure costs:

  • Corruption — data is destroyed. A chunk holds an unpaired surrogate and cannot be encoded, or input characters vanished.
  • Mangling — the bytes survive but the boundaries are wrong, so the text renders and embeds as nonsense. A cut inside a grapheme cluster.
  • Contract — a promise to the next stage is broken. An empty chunk, which embedding APIs reject; a chunk over the size limit.

Correctness of the checker

A checker that is wrong about Unicode is worse than none. Grapheme segmentation is a full implementation of UAX #29 extended grapheme clusters, including the rules for Indic conjuncts and emoji ZWJ sequences. It passes all 766 cases in Unicode's official GraphemeBreakTest.txt. The property tables are generated from the published Unicode Character Database by a script, not copied by hand, so a new Unicode version is a regeneration.

Decisions

Exit 0 is unreachable without a real cut. A splitter that returns the input as one chunk never placed a boundary, so nothing was tested. That exits 2, "nothing could be audited", not 0. Otherwise CI would go green on a splitter whose separator simply never appeared.

Normalisation is not destruction. A splitter that emits NFC text changed code points without losing anything. It is reported as mangling, not as data loss.

Overlap is not loss. Splitters that overlap chunks are handled so the overlap is never reported as missing or duplicated content.

What it found

Run against LangChain's RecursiveCharacterTextSplitter at a chunk size of 50, ten of the twelve cases fail grapheme integrity. The Rust-backed semantic-text-splitter passes all of them. That gap is the whole argument for running a check like this before choosing a splitter.

Honest limits

A pass is not a proof: splitlint tests the offsets it constructs, not every possible input. Grapheme integrity is graded as mangling rather than corruption because a splitter forced into a tiny chunk size may have no legal cut available.

Outcome

Point it at any importable splitter and it reports which invariants break, with the code points on both sides of every bad cut. Exit codes are designed so CI cannot go green on a splitter that was never actually tested.

loading index…Full retrieval trace →