Hybrid search and reranking: building a RAG pipeline that finds the right contract number or product code
Embeddings are good at understanding what a question means, but they easily mistake HD-2024-0153 for HD-2024-0135. This guide fixes that in six steps, from an eval set and a tokenizer through to reranking, and shows how to measure the result.
In brief
- Vector search matches on meaning, so it often misses exact codes. BM25 matches on character strings, so it catches them. You need both.
- RRF merges the two lists by rank, so there are no weights to tune. Reranking then filters again so only the most relevant chunks reach the model.
- On Anthropic's data, hybrid search cut retrieval failures by 49%, and adding reranking cut them by 67%. For your client, measure again on your own eval set.
Picture a member of a legal team typing into an internal chatbot: “What is the penalty clause in contract HD-2024-0153?” The RAG system answers fluently, but the answer comes from contract HD-2024-0135. Semantically the two strings are almost identical, and an embedding has no reason to tell them apart.
This kind of error turns up readily in document collections dense with identifiers, and Anthropic gives an example of the same kind: a user searching for “Error code TS-999” in a technical support database.
According to Anthropic, BM25 is particularly effective for queries containing unique identifiers or technical terms, because it looks for that exact text string in the documents.
On Anthropic’s dataset, combining embeddings with BM25 cut retrieval failures in the top 20 chunks by 49%, and adding reranking cut them by 67%. This guide walks through building that pipeline yourself and measuring the equivalent numbers on your own data.
What you will build, and what you need
The pipeline has five stages: the query runs in parallel through a keyword retriever and a vector retriever, the two lists are merged with reciprocal rank fusion (RRF), the result is reranked, and only then does it reach the LLM. You need Python 3, a set of chunks in the form {chunk_id: text}, and access to an embedding and reranking service. The examples below refer to Cohere.
The keyword search code in this article has been simplified for teaching purposes. In production, use the built-in BM25 in Elasticsearch or Weaviate.
Step 1: Build the eval before writing any code
Without measurements, you cannot show a client that you have improved anything. Collect real questions that contain codes, link each to its correct chunk, and measure how often the correct chunk fails to make the top 20. This is also the metric Anthropic uses.
def failure_rate(evalset, retrieve, top_n=20):
miss = sum(1 for q, gold in evalset
if gold not in retrieve(q)[:top_n])
return miss / len(evalset)
Check: run this function against your current vector pipeline and record the number. That is your baseline.
Step 2: The tokenizer must keep codes intact
When keyword search misses an identifier, a common cause is the tokenizer, not the algorithm. If “HD-2024-0153” is split into “hd”, “2024” and “0153”, the token “2024” will match hundreds of other contracts and becomes useless.
import re
TOKEN = re.compile(r"\w+(?:[-/]\w+)*")
def tokenize(text):
return [t.lower() for t in TOKEN.findall(text)]
tokenize("Hợp đồng HD-2024-0153 ký ngày")
# ['hợp', 'đồng', 'hd-2024-0153', 'ký', 'ngày']
The example input is Vietnamese (“contract HD-2024-0153 signed on”); note that the contract number survives as a single token.
Check: print the tokenized output for 20 real codes taken from the client’s data. Each code must come out as exactly one token.
Step 3: The keyword retriever (simplified)
The function below just counts how many times the query’s tokens appear in each chunk. It has none of the IDF or length normalisation of real BM25, but it is enough to show how string matching catches codes.
from collections import Counter
def keyword_search(query, chunks, top_n=20):
q = set(tokenize(query))
scored = []
for cid, text in chunks.items():
tf = Counter(tokenize(text))
score = sum(tf[t] for t in q)
if score:
scored.append((score, cid))
scored.sort(reverse=True)
return [cid for _, cid in scored[:top_n]]
Check: for a query containing HD-2024-0153, that contract’s chunk must rank above the chunk for HD-2024-0135. The HD-2024-0135 chunk may still appear in the list, because this simplified function also scores shared words such as “clause”, “penalty” and “contract”. If the two chunks tie, go back and check the tokenizer from step 2.
Step 4: The vector retriever, and a parameter that often gets forgotten
Elastic distinguishes between two kinds of search: keyword search matches on words, while semantic search matches on the meaning of the question. Vector search is still needed for questions such as “which contract has the heaviest penalty for late delivery”.
A common mistake sits in the embedding step: Cohere’s documentation requires queries to be embedded with input_type="search_query", while documents must use a different input_type value reserved for documents.
# giản lược: embed() là wrapper bạn tự viết quanh SDK
q_vec = embed([query], input_type="search_query")
(The comment reads: “simplified: embed() is a wrapper you write yourself around the SDK”.)
Check: grep the whole codebase to make sure the indexing step and the query step do not share the same input_type value.
Step 5: Merge with RRF, no weight tuning needed
Anthropic’s step is to merge the embedding and BM25 results using rank fusion and deduplicate them. The Elasticsearch documentation gives the RRF formula as follows: for each list, a document’s score is increased by 1/(k + rank). The advantage is that you do not have to find linear combination weights between the two retrievers.
def rrf(result_lists, k):
scores = {}
for results in result_lists:
for rank, cid in enumerate(results, start=1):
scores[cid] = scores.get(cid, 0.0) + 1.0 / (k + rank)
return sorted(scores, key=scores.get, reverse=True)
The scores dictionary handles deduplication on its own. You can set k to whatever value your engine’s documentation recommends.
A worked example by hand, with k = 10 chosen only to keep the numbers readable. Chunk A ranks 1st on the keyword side and 8th on the vector side: 1/11 + 1/18 ≈ 0.146. Chunk B ranks 1st on the vector side only: 1/11 ≈ 0.091. Chunk C ranks 2nd on both sides: 1/12 + 1/12 ≈ 0.167.
C wins despite topping neither list. That is the essence of RRF: it rewards chunks that both retrievers agree on.
If the client uses Weaviate, you do not need to write this function yourself. Weaviate has two fusion algorithms, of which rankedFusion keeps only each result’s position in each list and discards the scores. The alpha parameter sets the balance between dense and sparse: 0.5 is an even split, and the default is 0.75.
For query sets heavy with codes, try several alpha values on the eval set from step 1 rather than leaving the default.
Step 6: Rerank, with YAML formatting for structured data
According to Anthropic, reranking is a filtering technique used to ensure that only the most relevant chunks are passed to the model. Take the top chunks after RRF, send them to Cohere Rerank along with the query, and keep only the first few. Cohere’s documentation recommends that if documents contain structured data, they should be formatted as YAML strings for best performance.
# giản lược: không xử lý escape ký tự đặc biệt
def to_yaml(rec):
return "\n".join(f"{k}: {v}" for k, v in rec.items())
to_yaml({"so_hop_dong": "HD-2024-0153",
"ben_ban": "...", "dieu_khoan_phat": "..."})
(The comment reads: “simplified: does not escape special characters”. The keys mean contract number, seller and penalty clause.)
So for contract records pulled from an ERP, follow that recommendation: convert them into one key: value line per field before reranking, rather than sending raw JSON.
Check: rerun failure_rate for three configurations (vector, hybrid, hybrid + rerank) and put the three numbers side by side in a table.
Mistakes that make a pipeline look right but behave wrongly
The first is a tokenizer that splits codes, as described in step 2. It raises no exception; results simply get steadily worse. The second is indexing documents with the input_type meant for queries, usually because code was copied over from an experimental notebook. The third is sending structured records to the reranker while ignoring the YAML formatting recommendation.
The fourth is in how results are reported. The 49% and 67% figures are Anthropic’s results on Anthropic’s data, not a promise for your client’s data. Do not put them on a slide as if they were your results. Present the measurements from step 1.
How this skill shows up at a client site
In your first week with a client, the first thing worth asking is: do users often type product codes, contract numbers or error codes, and what format do those codes follow? The answer determines how you write the tokenizer, before any discussion of which embedding model to pick.
When reading job descriptions for FDE or AI engineer roles, phrases such as “retrieval quality”, “hybrid search” and “evaluation” refer to exactly this skill. On your CV, avoid vague lines like “built RAG”.
Write it in this form instead: “reduced top-20 retrieval failure rate from X% to Y% on a 200-question eval set containing contract codes, using hybrid search and reranking”, where X and Y are numbers you measured yourself.
The client will not remember whether you used RRF or what alpha you chose. They will remember the first time the chatbot answered about the right contract, HD-2024-0153.