Muzamil Chaudhery Writing
Building Rasikh

Why Rasikh's search has no relevance cutoff

· 5 min read

Rasikh is a free, Urdu-first Islamic AI for the Quran and hadith. You can ask it a question in Urdu, Roman Urdu or English and it finds the verses and narrations that answer it. Under the hood, that starts with search: turning a question like “safai nisf iman” into the right handful of texts.

Early on I wanted Rasikh to do something that sounds obvious: when a question has nothing to do with the Quran or hadith, say so instead of returning the “closest” results anyway. The standard way to do that is a relevance cutoff. Every result gets a similarity score, and anything below some threshold gets dropped.

I measured it properly before shipping it. It didn't work.

The numbers that killed the idea

Rasikh's search uses a multilingual embedding model (bge-m3), so every question and every text becomes a vector, and “how relevant” becomes “how close are these vectors” (cosine similarity). I ran a set of real questions and a set of deliberately off-topic ones, and looked at the best score each one got.

QuestionKindBest score
aaj ka mausam kaisa hai (what's the weather today)Off-topic0.61
safai nisf iman (cleanliness is half of faith)Real0.45

A question about the weather scored higher than a question about one of the best-known hadith. That wasn't a fluke. Across the whole set, real questions scored between 0.44 and 0.86, and off-topic ones between 0.39 and 0.61. The ranges overlap. Any threshold low enough to keep the real questions lets the weather through; any threshold high enough to block the weather throws away real answers.

A similarity score tells you which text is closest. It doesn't tell you whether anything is close enough.

So I made the decision to have no cutoff at all. Search always returns its best matches, and it's good at that part. Deciding whether a question belongs in Rasikh at all became a separate job: a small intent step that reads the question and classifies it before search runs. The off-topic questions I'd collected became its test set.

You can't improve what you don't measure

The only reason I caught this is that I'd built an evaluation set first, by hand. For a long list of real questions I went through the results and marked which ones were actually relevant: 1,280 judgments in the current baseline.

Two details made that set trustworthy:

Against that set, the current search gets the right text into the top ten for every question (Hit@10 of 1.0), usually puts a relevant result first (MRR@10 of 0.919), and finds about three-quarters of all relevant texts in the top ten (recall@10 of 0.744). That last number is the one I'm working on.

What changed because of it

The lesson I keep relearning

The obvious fix was a one-line change: if score < threshold. It would have looked like it worked in a quick demo, and quietly hidden good answers from real people. Measuring first turned a bad one-liner into a better design. That's true well beyond search, and it's how I try to build everything now.

← All writing