Why Rasikh's search has no relevance cutoff
Rasikh is a free, Urdu-first Islamic AI for the Quran and hadith. You can ask it a question in Urdu, Roman Urdu or English and it finds the verses and narrations that answer it. Under the hood, that starts with search: turning a question like “safai nisf iman” into the right handful of texts.
Early on I wanted Rasikh to do something that sounds obvious: when a question has nothing to do with the Quran or hadith, say so instead of returning the “closest” results anyway. The standard way to do that is a relevance cutoff. Every result gets a similarity score, and anything below some threshold gets dropped.
I measured it properly before shipping it. It didn't work.
The numbers that killed the idea
Rasikh's search uses a multilingual embedding model (bge-m3), so every question and every text becomes a vector, and “how relevant” becomes “how close are these vectors” (cosine similarity). I ran a set of real questions and a set of deliberately off-topic ones, and looked at the best score each one got.
| Question | Kind | Best score |
|---|---|---|
| aaj ka mausam kaisa hai (what's the weather today) | Off-topic | 0.61 |
| safai nisf iman (cleanliness is half of faith) | Real | 0.45 |
A question about the weather scored higher than a question about one of the best-known hadith. That wasn't a fluke. Across the whole set, real questions scored between 0.44 and 0.86, and off-topic ones between 0.39 and 0.61. The ranges overlap. Any threshold low enough to keep the real questions lets the weather through; any threshold high enough to block the weather throws away real answers.
A similarity score tells you which text is closest. It doesn't tell you whether anything is close enough.
So I made the decision to have no cutoff at all. Search always returns its best matches, and it's good at that part. Deciding whether a question belongs in Rasikh at all became a separate job: a small intent step that reads the question and classifies it before search runs. The off-topic questions I'd collected became its test set.
You can't improve what you don't measure
The only reason I caught this is that I'd built an evaluation set first, by hand. For a long list of real questions I went through the results and marked which ones were actually relevant: 1,280 judgments in the current baseline.
Two details made that set trustworthy:
- Question groups. People ask the same thing in different words and scripts: Urdu, Roman Urdu, English, sometimes with Arabic terms. Questions that mean the same thing are grouped, so a judgment on one counts for its siblings.
- Only human labels count. Rasikh has curated topics that help ranking, but labels derived from those topics are never used for scoring. Otherwise the system would be grading its own homework.
Against that set, the current search gets the right text into the top ten for every question (Hit@10 of 1.0), usually puts a relevant result first (MRR@10 of 0.919), and finds about three-quarters of all relevant texts in the top ten (recall@10 of 0.744). That last number is the one I'm working on.
What changed because of it
- Search is deterministic. The same question always gives the same results in the same order, so a score change means the system changed, not luck.
- Curated topics, with Arabic terms. 42 hand-built topics help connect everyday phrasing to the vocabulary the texts actually use.
- Every reference is checked against the stored text. If something can't be verified, Rasikh shouldn't guess.
The lesson I keep relearning
The obvious fix was a one-line change: if score < threshold. It would have looked like it worked in a quick demo, and quietly hidden good answers from real people. Measuring first turned a bad one-liner into a better design. That's true well beyond search, and it's how I try to build everything now.