Journal · SQLite
A longer prefix returned fewer results
Search-as-you-type was going blank mid-word and then recovering. The cause turned out to be a property of the Porter stemmer that makes prefix matching non-monotonic — and it's reproducible in four lines of SQL.
The symptom
Typing a query one character at a time should narrow results monotonically. Every additional character can only ever remove matches, never add them back.
That isn't what happened. Typing animals into our search box produced results at anim, nothing at all at anima, and results again at animal. The grid emptied out for exactly one keystroke and then repopulated.
The index is SQLite FTS5, declared like this:
CREATE VIRTUAL TABLE search_index USING fts5(
content_item_id UNINDEXED,
raw_text,
preview_title,
tags,
tokenize = 'porter unicode61 remove_diacritics 1'
);
Queries are issued as prefix terms — anim*, anima*, and so on — as the user types.
Minimal reproduction
No application code needed. This is plain SQLite — verified on 3.49.1.
CREATE VIRTUAL TABLE t USING fts5(body, tokenize='porter unicode61');
INSERT INTO t VALUES ('cute animals playing in the snow');
SELECT count(*) FROM t WHERE t MATCH 'anim*'; -- 1
SELECT count(*) FROM t WHERE t MATCH 'anima*'; -- 0 <-- ?
SELECT count(*) FROM t WHERE t MATCH 'animal*'; -- 1
Swap the tokenizer for unicode61 alone and all three return 1, as you'd expect.
Why it happens
The porter tokenizer stems both the stored tokens and the query terms. "animals" and "animal" are both stored as the stem anim.
A prefix query is stemmed first, then prefix-matched against the stored stems. So the real question for any typed prefix P is: is stem(P) a prefix of anim?
| Typed | Stems to | Prefix of anim? | Result |
|---|---|---|---|
anim* | anim | yes | match |
anima* | anima — no rule applies | no, it's longer | zero |
animal* | anim | yes | match |
That's the whole thing. Stemming shortens tokens, so an intermediate prefix of a word can be longer than the stem that word is actually stored under. When no stemming rule happens to apply to that partial string, it stays long, overshoots the stored stem, and matches nothing.
Prefix matching assumes the stored form is a superstring of what you type. Stemming breaks that assumption, and the two features don't compose.
It isn't one unlucky word
"animals" was just what we happened to test with. The dead zone shows up across ordinary English, and it's often wider than one character. Each row below is a single word, indexed alone, queried with every prefix from three characters up:
| Word | Prefix results (1 = found, 0 = nothing) |
|---|---|
animals | ani1 anim1 anima0 animal1 animals1 |
cooking | coo1 cook1 cooki0 cookin0 cooking1 |
running | run1 runn0 runni0 runnin0 running1 |
swimming | swi1 swim1 swimm0 swimmi0 swimmin0 swimming1 |
training | tra1 trai1 train1 traini0 trainin0 training1 |
happiness | hap1 happ1 happi1 happin0 happine0 happines0 happiness1 |
national | nat1 nati1 natio1 nation1 nationa0 national1 |
traveling | tra1 trav1 trave1 travel1 traveli1 travelin0 traveling1 |
Every one of these has a hole. happiness and swimming each go dark for three consecutive keystrokes. And note running: the stem is run, so even runn — a four-character prefix of a seven-character word — finds nothing.
The -ing class is the painful one, because it's a large slice of the vocabulary people actually search for.
Why this is worse than a normal empty result
A search that returns nothing is fine when there's genuinely nothing. The problem here is that it looks like a bug in the app rather than a fact about the data — the results were right there a keystroke ago.
It also breaks incremental search specifically. If results only shrink as you type, you can cache, you can filter the previous result set, and the user builds an accurate mental model. Once an extra character can take you from 40 results to zero and back to 40, every one of those assumptions is gone.
The interim fix, and why it's a workaround
The proper fix is to drop the stemmer. We didn't do that first, because changing the tokenizer means dropping and rebuilding the index on every installed device — a migration with real failure modes, and this was a shipped app with a user-visible bug.
So the interim fix is a backoff. On a zero result, walk the last token back one character at a time and use the first form that matches:
anima* -> 0 results
anim* -> match use this
Three details that matter more than the idea:
- There's a three-character floor. Below that the backoff would match most of the index and the results would be noise.
- If nothing matches at any length, the original expression is restored — so a genuinely empty search still honestly reports no matches, rather than silently showing results for a shorter word the user never typed.
- The result grid and the "N saved" counter resolve through the same function. An earlier version computed them separately, which meant the counter could disagree with the number of visible tiles.
It's a workaround and it reads like one. It papers over the hole rather than removing it, and it can return results for anim when the user typed anima — technically not what they asked for, though in practice indistinguishable from what they wanted.
The actual fix
Drop porter and use unicode61 alone.
The reason this is acceptable rather than a loss is that prefix matching already provides most of what the stemmer was there for. The classic argument for stemming is that a search for "animal" should find "animals" — but animal* finds "animals" anyway, without a stemmer, because it's a literal prefix. Stemming earns its keep on suffixes that aren't prefixes, like "ran" for "run", and for a search box over your own saved content that's a much smaller win than the dead zone is a loss.
The cost is the migration: dropping and rebuilding the FTS index on device, for every install, without losing anything if it's interrupted. That ships in the next release rather than being rushed out behind a bug fix.
The general lesson
Prefix search and stemming are individually reasonable and jointly broken. If you're building search-as-you-type on FTS5, or on any engine that stems at index time, the property you actually depend on is that results shrink monotonically as the query grows — and a stemmer silently removes it.
It's easy to miss because it doesn't fail on the words you test with. It fails in the middle of words, on a subset of the vocabulary, and only if you happen to type one character at a time and look.
If you have a stemmed FTS5 index behind an incremental search box right now, the four-line reproduction at the top will tell you in about a minute whether you have this.
Context
This came out of MessageVault AI, an iPhone app that saves posts you share from Instagram, Facebook, YouTube, X and TikTok, tags them with an AI model at save time, and then lets you search them.
The relevant architectural point is that search itself invokes no model and touches no network — the AI runs once, when a post is saved, and writes tags into the local index. Everything after that is a SQLite query on the device. That's why an FTS5 tokenizer setting was able to become a user-visible bug in the first place, and it's described in more detail here.