- Kind
- Own open source · MIT
- Year
- 2026
- Role
- Author
- Stack
- PythonUnicode UAXpytestGitHub Actions
- Source
- r0h1tb/splitlint ↗
splitlint
A CLI and library that audits any RAG text splitter for the Unicode corruption a round-trip test cannot see.
The hard partThe broken chunk still joins back into the original string. Nothing throws; retrieval just gets worse for one language.
- 766 of 766 official Unicode grapheme conformance cases pass
- LangChain's RecursiveCharacterTextSplitter fails 10 of its 12 cases
- Grew out of an upstream fix in langchain4j
The problem
Every retrieval pipeline chunks text before embedding it, and almost every chunker cuts every N "characters". The bug is in what counts as a character.
Count UTF-16 code units, as Java's substring and JavaScript's slice
do, and a cut can land between the two halves of a surrogate pair. That
chunk can no longer be encoded as UTF-8 at all. Count code points, as
Python does, and the cut lands inside a family emoji, a Devanagari
conjunct or a Hangul syllable instead. Either way the chunks still join
back into the original text, so the obvious test passes. The damaged
chunk embeds as something else, and you find out when retrieval
quality drops for one script.
I found the first version of this in langchain4j: its splitter worked in UTF-16 code units, and the fix (#6308, in review) splits by code point. splitlint is that investigation turned into a tool that checks any splitter in seconds.
Try it
The same three ways of counting, on text a reader would call ordinary. Change the chunk size and watch where each one cuts.
Splitter · same text, three ways to count
This demo runs in your browser. With JavaScript off, the case study below describes the same three failure modes.
Every row joins back into the original text, so a round-trip test passes all three. The broken chunks embed as something else. Grapheme boundaries here come from your browser’s Intl.Segmenter; splitlint ships its own UAX #29 implementation, checked against all 766 official test cases.
How it finds the bug
A family emoji in the middle of a chunk proves nothing. Each of the twelve test cases pads the text with separator-free filler until the hazard straddles the splitter's cut offset exactly. That also pushes recursive splitters past their paragraph, line and space separators, down to the character-level fallback where the boundary bugs live.
Six invariants, graded by what the failure costs:
- Corruption — data is destroyed. A chunk holds an unpaired surrogate and cannot be encoded, or input characters vanished.
- Mangling — the bytes survive but the boundaries are wrong, so the text renders and embeds as nonsense. A cut inside a grapheme cluster.
- Contract — a promise to the next stage is broken. An empty chunk, which embedding APIs reject; a chunk over the size limit.
Correctness of the checker
A checker that is wrong about Unicode is worse than none. Grapheme
segmentation is a full implementation of UAX #29 extended grapheme
clusters, including the rules for Indic conjuncts and emoji ZWJ
sequences. It passes all 766 cases in Unicode's official
GraphemeBreakTest.txt. The property tables are generated from the
published Unicode Character Database by a script, not copied by hand,
so a new Unicode version is a regeneration.
Decisions
Exit 0 is unreachable without a real cut. A splitter that returns the input as one chunk never placed a boundary, so nothing was tested. That exits 2, "nothing could be audited", not 0. Otherwise CI would go green on a splitter whose separator simply never appeared.
Normalisation is not destruction. A splitter that emits NFC text changed code points without losing anything. It is reported as mangling, not as data loss.
Overlap is not loss. Splitters that overlap chunks are handled so the overlap is never reported as missing or duplicated content.
What it found
Run against LangChain's RecursiveCharacterTextSplitter at a chunk size
of 50, ten of the twelve cases fail grapheme integrity. The Rust-backed
semantic-text-splitter passes all of them. That gap is the whole
argument for running a check like this before choosing a splitter.
Honest limits
A pass is not a proof: splitlint tests the offsets it constructs, not every possible input. Grapheme integrity is graded as mangling rather than corruption because a splitter forced into a tiny chunk size may have no legal cut available.
Outcome
Point it at any importable splitter and it reports which invariants break, with the code points on both sides of every bad cut. Exit codes are designed so CI cannot go green on a splitter that was never actually tested.