Hybrid search in Postgres: BM25, pgvector, RRF and a reranker
No separate search engine. One Postgres database, two rankings, one fusion and a reranker — with the parameters we actually run.
Vector search finds text that *means* the same thing as your question. It is bad at finding text that contains the exact string you typed: an invoice number, a surname, a library version. Keyword search has the opposite profile. Hybrid search runs both and merges the results, and for an AI memory — where people ask both kinds of question — it is the difference between "it remembered" and "it found something vaguely related".
If you want the concepts first, our explainer on how persistent AI memory works covers them without code. This article is the implementation: the pipeline Pryzm Memory runs in production, inside a single PostgreSQL database, with the SQL, the fusion code and the settings we chose — and the mistakes that made us choose them.
The pipeline at a glance
- Keyword arm. PostgreSQL full-text search with a BM25-style score, 50 candidates.
- Vector arm. pgvector with an HNSW index on multilingual embeddings, at least 50 candidates.
- Fusion. Reciprocal rank fusion (k = 60) merges the two ranked lists.
- Reranking. A multilingual cross-encoder rescores the top 40.
- Freshness. A small, decaying bonus for recent memories, applied after the reranker.
Each step is cheap to add to an existing Postgres setup. Most of the work is in the details below.
The keyword arm: BM25-style scoring in SQL
PostgreSQL has full-text search built in, but its ranking functions are not BM25: they do not weigh a term by how rare it is across your documents. Rarity is exactly what makes keyword search useful next to vectors — a rare identifier should count far more than a common word. So we compute the inverse document frequency (IDF) ourselves:
-- Bras mots-clés : score de type BM25 (IDF, tf binaire), en SQL pur.
WITH terms AS (
SELECT DISTINCT quote_literal(t)::tsquery AS tq
FROM unnest(tsvector_to_array(
to_tsvector('french', unaccent($1)))) AS t
),
docs AS MATERIALIZED (
SELECT id, to_tsvector('french', unaccent(content)) AS v
FROM memories
),
total AS (SELECT count(*)::float8 AS n FROM docs),
weights AS (
SELECT terms.tq,
ln(1 + (total.n - df.df + 0.5) / (df.df + 0.5)) AS idf
FROM terms, total,
LATERAL (SELECT count(*)::float8 AS df
FROM docs WHERE docs.v @@ terms.tq) df
WHERE df.df > 0
)
SELECT d.id, sum(w.idf) AS score
FROM docs d JOIN weights w ON d.v @@ w.tq
GROUP BY d.id
ORDER BY score DESC
LIMIT 50;Three choices in that query came from bugs, not from theory:
- OR, not AND. Our first version used websearch_to_tsquery, which requires every word of the query. A natural question ("what did we decide about the pricing page?") almost never contains only words present in one memory, so the keyword arm returned nothing. Splitting the query into terms and summing their weights fixed it.
- IDF per user. Each account's memories are scored against that account's own corpus. A word that is rare in your notes is informative for you, whatever it is in someone else's.
- Accents removed on both sides. We indexed text with accents and searched without them, so French queries silently missed. unaccent has to be applied to the documents and to the query, with the same language configuration (here, French stemming).
Term frequency is binary here (a term is present or not), which makes this a simplified BM25. For short memories that loses little. In production, store the tsvector in an indexed column instead of computing it per query.
The vector arm: pgvector, HNSW and filters
Embeddings come from multilingual-e5-small (384 dimensions). The e5 family expects a prefix — "query: " for the question, "passage: " for stored text — and quietly performs worse without it. Vectors live in pgvector behind an HNSW index, with ef_search set to 100.
The trap is filtering. An approximate index is scanned first and filtered after, so a search restricted to one project can come back with far fewer rows than you asked for. pgvector 0.8 added iterative index scans, which keep scanning until enough rows pass the filter. We run them in relaxed_order mode.
Reciprocal rank fusion: merging scores that share no units
A cosine distance and a keyword score cannot be added: they are not on the same scale, and the keyword score changes with the size of the corpus. Reciprocal rank fusion ignores the raw scores and uses only positions. Each document gets 1 / (k + rank) for each list it appears in, and the sums are sorted:
def rrf(dense_ids, lexical_ids, k=60):
scores = {}
for ranking in (dense_ids, lexical_ids):
for rank, doc_id in enumerate(ranking, start=1):
scores[doc_id] = scores.get(doc_id, 0.0) + 1.0 / (k + rank)
return sorted(scores, key=scores.get, reverse=True)k = 60 is the value from the original paper and we kept it: it flattens the gap between rank 1 and rank 5, so a document that is decent in both lists beats one that is first in a single list. Ask each arm for more candidates than you need (we take 50 from each) — fusion only helps when the two lists overlap a little.
The reranker: where most of the quality comes from
The first two arms compare a question and a document that were each turned into numbers separately. A cross-encoder reads the two together and scores how well this document answers this question. It is much more accurate and much slower, so it only sees the shortlist:
from sentence_transformers import CrossEncoder
reranker = CrossEncoder("cross-encoder/mmarco-mMiniLMv2-L12-H384-v1",
max_length=128, device="cpu")
def rerank(query, fused, n=40):
head, tail = fused[:n], fused[n:]
scores = reranker.predict([(query, doc["content"]) for doc in head])
for doc, s in zip(head, scores):
doc["score"] = float(s)
return sorted(head, key=lambda d: d["score"], reverse=True) + tailWe use a multilingual cross-encoder because our users write in French, German, Spanish and more; many rerankers are trained on English only and degrade elsewhere without warning. Input is cut at 128 tokens, which is enough for a memory and keeps the cost down. On a CPU server, the reranker is by far the slowest step — size the shortlist with that in mind.
Dates: what the reranker cannot see
A memory store accumulates versions of the same fact: "the launch is on the 3rd", then "the launch moved to the 5th". The reranker judges relevance and ignores dates, so both versions score alike and the old one can win. We add a decaying freshness bonus *after* reranking, with a half-life of 180 days:
def with_recency(score, age_days, weight, half_life=180):
return score + weight * 2 ** (-age_days / half_life)Applied after the reranker, it only breaks near-ties. Applied before, it would let a recent but irrelevant memory push out the right answer.
When the keyword arm fails
The keyword query runs inside a SAVEPOINT. If it fails — a missing extension, an unsupported language configuration — we roll back to the savepoint, log it, and continue with vector results only. A degraded search is better than an error in the middle of someone's conversation.
When you do not need all this
- A few hundred documents, one language, mostly conceptual questions: pgvector alone is fine.
- Tight latency budget on CPU: start with fusion, add the reranker when you can measure what it changes.
- Very large corpora or heavy faceting: a dedicated search engine will serve you better than SQL.
For an AI memory, where the same person asks "what was the invoice number for client X?" and "what did we think about pricing?" in the same afternoon, we found every stage worth keeping. If you would rather use it than build it, that pipeline is what answers when Claude or ChatGPT queries Pryzm Memory — see how MCP memory servers compare.
Frequently asked questions
What is hybrid search?
Running a keyword search and a vector (semantic) search on the same documents, then merging the two rankings, usually with reciprocal rank fusion. It finds both exact strings and paraphrases.
Does PostgreSQL support BM25?
Not natively: its full-text ranking functions do not use inverse document frequency. You can compute a BM25-style score in SQL, as shown above, or use an extension.
What value of k should I use for reciprocal rank fusion?
60, the value from the original paper, is a sound default. Lower values reward the top of each list more; higher values reward agreement between lists.
Is a reranker worth the extra latency?
It is usually the step that improves results the most, but it is also the slowest. Run it only on a shortlist (we use 40 candidates) and measure on your own data.