Back to Blog

Evoke for Postgres: Meaning, indexed like keywords

Search PostgreSQL by meaning, with no embedding pipeline to run. Keyword and semantic evidence share one index and one score, and a 30M-parameter model runs on CPUs inside the database.

DOCUMENTSONE INDEXONE RANKED LISTkeywordmeaning12345
keyword evidence semantic evidence
Move through the stream, or click to send a burst.

No AI system, however capable, can use a document it never found. Whatever an agent is asked to do, it works from what its search step returned. Agents also search often: they look something up, read the result, sharpen the question and search again, several times in a single task.

For teams whose data lives in PostgreSQL, adding search by meaning has usually meant adding machinery. There is an embedding model or API to call, a job to keep vectors in step with the rows, a vector index beside the keyword index, and a query that merges two ranked lists into one.

Today we're releasing Evoke, a PostgreSQL extension that replaces that machinery with a single index and a compact model running inside the database. The code and the model are open source under Apache-2.0.

Evoke in 30 seconds: one index for keywords and meaning, inside Postgres.

Keywords and meaning in one index

Keyword search scored with BM25 is good at the literal: product codes, case numbers, error strings and names. It misses documents that make the same point in different words. Semantic search finds those, but in the common Postgres setup it runs in a separate vector index. Its scores sit on a different scale, so the two result lists have to be merged afterwards.

Evoke uses learned-sparse retrieval instead. Its encoder is built on IBM's Granite-Embedding-30M-Sparse, a model of about 30.3M parameters. It turns each document and each query into weighted terms from a fixed vocabulary. Those terms can include related words the text never uses. The semantic terms go into the same inverted index as the BM25 keyword terms, in a namespace of their own. At query time, both kinds of evidence add up to one score per document.

Try it · illustrative
Meaning, written as terms the index already understands
Query
“heart attack warning signs”
Keyword termsheartattackwarningsigns
Meaning termsmyocardialinfarctionsymptomschestpaincardiac
Ranked results
2
Recognising myocardial infarction2.73
Chest pain spreading to the arm, breathlessness and sweating are typical presenting symptoms.
1
Heart attack: what to do2.99
Call emergency services at once if you suspect a heart attack.
3
Warning signs on site2.00
Hazard signs and warning labels for construction areas.
4
Cardiac rehabilitation0.43
Exercise and education after a cardiac event.
keyword evidence semantic evidence
Illustrative: the terms and weights are hand-set to show the mechanism, not taken from the model. In Evoke both kinds of evidence sit in one inverted index, each in its own namespace, and add up to one score per document. Try “Exact codes”: the literal match still leads, because keyword evidence is never given up.
The query path
One lookup replaces two searches and a merge
Common setup in Postgres today
Text query
App calls an embedding model
outside the database
Keyword search + vector search
two indexes
Merge the two rankings
two score scales
Results
5 steps2 indexes1 merge
Evoke
Text query
One index lookup
Keyword and semantic evidence, one score. Model runs inside Postgres.
one index
Results
3 steps1 index0 merges

Because the semantic signal is expressed as vocabulary terms, it fits the structure a keyword index already has. Evoke needs no dense vector index, and each query returns a single ranked list, so there is nothing to fuse. Applications work through ordinary SQL, with CREATE INDEX ... USING ii42 to build the index and the ii42_query(...) functions to return ranked rows or explicit hits.

The database runs the model

The official packages and the Docker image bundle the model with ONNX Runtime, so there is nothing separate to download. It runs on ordinary CPUs in shared background workers, and each database connection hands its encoding work to those workers instead of loading its own copy of the model. No GPU is needed, and no text leaves the database to be embedded.

When a row is inserted or updated, the write stores its keyword evidence straight away and queues the row for semantic encoding. The workers encode it afterwards, so the write never waits for the model.

The write path
The write commits now. Meaning follows.
Your table
No rows yet.
One index
keyword namespace
meaning namespace
Shared workers
Worker 1 · idle
Worker 2 · idle
Queue
empty
Click INSERT a few times. Each write stores its keyword evidence and commits straight away; the row joins a queue, and background workers that every connection shares encode it with the model afterwards. No GPU, and no text leaves the database.

Queries take plain text. The application sends the question, and the database encodes it and searches. Every returned row is checked against the table as it currently stands, so a deleted document cannot resurface from a stale copy. Index maintenance follows PostgreSQL's own rules for row visibility, crash recovery and physical replication.

Those choices are what make the design lean. There is one index where the common setup has two, one model runtime shared by every connection, and no pipeline outside the database to keep in step with the data.

What the numbers show

We evaluated the P2.1 model on two fixed English suites, BEIR15 and MTEB10, against BM25 and against a 0.6B-parameter dense embedding model, perplexity-ai/pplx-embed-v1-0.6B, searched through VectorChord. The release ships the P2.2 model, built on the same Granite-based route.

Recall@100 is the share of relevant documents that reach the first 100 results, and CUB@1000 is the same share for a candidate pool of up to 1,000.

Results
Within reach of a model twenty times its size
BM25best
keyword only
0.563
Evokebest
P2.1 model · 30.3M parameters
0.667
0.6B densebest
pplx-embed-v1-0.6B via VectorChord
0.671
Model size
Evoke · 30.3M
dense · 0.6B
Recall@100 is the share of relevant documents that reach the first 100 results; CUB@1000 is the same share for a candidate pool of up to 1,000. Bars run from 0 to 1. Evaluated with the P2.1 model; the release ships P2.2, built on the same Granite-based route. Each square is about 30M parameters.

Against keyword search alone, Evoke raised Recall@100 from 0.563 to 0.667 on BEIR15 and from 0.595 to 0.703 on MTEB10. Against a dense model with roughly twenty times the parameters, it matched top-100 recall to within 0.004 and placed more relevant documents in the top 1,000, on both suites.

First-stage retrieval decides which documents the rest of the pipeline can work with, and a reranker can only reorder what it receives. Finding more of the relevant material at this stage raises the ceiling for everything that follows.

Where Evoke fits

Evoke suits English text in PostgreSQL, with results feeding a reranker, an agent or a person who reviews them. It is a natural upgrade for teams that already search Postgres by keyword and want matches by meaning without building a vector stack. For teams that already run keyword and vector search side by side, Evoke brings both into one index with one model.

Evoke also suits deployments where text should not leave the server, including air-gapped ones, since the release includes a checksummed Docker archive for offline loading. Good first projects include agent and RAG search over internal documentation and policies, and published collections such as clinical guidelines, case law or regulatory rulebooks. Review work is a strong fit too, wherever finding everything relevant counts for more than finding the single best match.

Try it

The quickest start is the PostgreSQL 18 image. On a fresh data volume it creates the extension and enables the shared runtime.

bash
docker pull ghcr.io/intelligent-internet/ii-42:pg18-v0.2.5

Start the container with the command in the README, then build an index and query it:

sql
CREATE INDEX docs_semantic_idx ON docs USING ii42 (body) WITH (sae = true);

SELECT d.id, d.title,
       ii42_query('docs_semantic_idx'::regclass, 'database search architecture') AS score
FROM docs AS d
ORDER BY score DESC
LIMIT 10;

BM25 is the default, and sae = true turns on Sparse Semantic Retrieval (SSR). The extension, image and SQL functions use the ii42 name.

Official packages for PostgreSQL 17 and 18 on Linux x86-64, with checksums and an offline Docker archive, are on the v0.2.5 release page. Package installs load the extension through shared_preload_libraries, followed by a restart, and existing databases follow the upgrade guide. Source builds take the model from Hugging Face.

Questions, issues and results from your own data are welcome on GitHub.

Credits

Evoke is built on Granite-Embedding-30M-Sparse from IBM's Granite Embedding team, released under Apache-2.0. On top of that encoder, Evoke adds compilation, calibration, posting publication, scoring and database integration.

Download Evoke · Get started · Model · Model report · System report