An open-source C++ search engine indexed 2.2M records in 3.6 minutes, then searched them…
@curiouslearner · LinkedIn post · curiouslearner_an-op… original · 2026-07-07
Post
An open-source C++ search engine indexed 2.2M records in 3.6 minutes, then searched them in 11ms, on 4 vCPUs.
But the speed isn’t the interesting part.
For decades, search has worked one way => you type keywords, the engine matches keywords, and typos are your problem. Misspell a word and you get zero results, even when the answer was sitting right there in the index.
Typesense, the engine behind that benchmark, treats typo tolerance as a default, not a setting you configure. A query with a spelling mistake still returns the right result . Sounds minor, until you count how often real users, thumbing a query on a phone, get a word slightly wrong.
The idea worth studying, though, is what Typesense calls Natural Language Search.
Nobody thinks the way a database expects. When you want a book, you don’t think “genre = mystery, country = Japan, price < 15.” You think, and type: “mystery novels set in Japan under fifteen dollars.”
That one sentence hides several constraints, expressed in ordinary language no database can execute directly. Most search systems either ignore that nuance or force a developer to hand-write parsing logic for it.
Typesense’s Natural Language Search hands the sentence to an LLM, which converts free-form phrases into the structured filters, sorts, and queries the engine actually runs . The user supplies intent in plain language. The system translates it into something executable. Then retrieval happens the way it always has, just with far better inputs.
This is the pattern I think we’ll see everywhere in applied AI => The LLM’s job often isn’t to produce the final answer, it’s to translate an ambiguous human request into a structured instruction a reliable, deterministic system can execute well.
The model handles understanding. The engine handles execution.
If you’re building search and your users type full sentences instead of keywords, this is a design pattern worth studying, regardless of which engine you land on.
GitHub repo in the comments. 👇