A trustworthy retrieval-augmented generation (RAG) system answers only from sources it can cite, and declines when it cannot find any. That contract is easy to state and hard to keep, because it has to hold across search quality, prompt design, multi-tenancy and testing. This guide shows how to build it on PostgreSQL with pgvector: HNSW vector indexes, hybrid search with reranking, and tenant isolation enforced inside the database with Row-Level Security.
The contract: cite or decline
State the rule precisely. Every answer must carry citations to the passages it used, and if retrieval returns nothing that supports an answer, the system says so instead of generating one.
Enforce it structurally, not by prompt alone. The retrieval layer returns passages with stable identifiers. The generation step is asked to answer only from those passages and to reference them. A check after generation rejects any answer whose citations do not match retrieved passages. An answer with no valid citation becomes a refusal.
Storing and searching: pgvector with HNSW
pgvector adds a vector type and similarity search to PostgreSQL. Keeping embeddings next to the relational data (documents, tenants, permissions) means one system to back up, secure and query.
For search speed, HNSW (Hierarchical Navigable Small World) indexes give approximate nearest-neighbour search with good recall at low latency. Their parameters trade recall against build time, memory and query cost, so tune them against your own evaluation queries rather than accepting the defaults.
Hybrid search and reranking
Vector search finds passages that mean something similar. Keyword search finds passages that contain the exact terms, such as product names, identifiers and error codes, that embeddings can blur. Hybrid search runs both and merges the results.
A reranker then re-scores the merged candidates with a model that reads the query and the passage together, which usually improves the top of the list that the generator actually sees. Reranking costs latency, so rerank a small candidate set, dozens of passages rather than hundreds.
Tenant isolation with Row-Level Security
In a multi-tenant system, the worst failure is one tenant seeing another tenant’s passages. Row-Level Security (RLS) in PostgreSQL enforces access rules in the database: policies attach to tables, and every query, including the vector search, only sees the rows that the session’s tenant may read.
Because the rule lives in the database, application code cannot skip it, provided the application connects with a role that neither owns the tables nor bypasses RLS (table owners are exempt from policies unless the table uses FORCE ROW LEVEL SECURITY). The guarantee is also testable at the data layer: run the same query under two tenant identities and assert that the result sets do not overlap.
An owned retrieval API
Put retrieval behind an API that the platform owns instead of letting each product query the database. The API is where the contract is enforced: it sets the tenant for the session, runs hybrid search and reranking, returns passages with identifiers, and applies the citation policy. Products get one retrieval contract, and improvements to search benefit all of them at once.
In practice
On the ENSO platform, I built the RAG architecture on Aurora PostgreSQL with pgvector (HNSW, hybrid search and reranking) behind an owned retrieval API with a strict citation policy: every answer cites its sources or declines, and tenant isolation is enforced in the database with Row-Level Security. The surrounding architecture is described in the ENSO case study.
Trade-offs and limits
Putting isolation and retrieval in PostgreSQL ties the design to one engine and to its scaling limits; a very large corpus may need a dedicated vector store.
Strict citation increases refusals. Users will see “I can’t answer that from the available sources” more often, which is the intended behaviour but needs a clear message and a path to add sources.
Reranking and hybrid search add latency and moving parts, and every parameter (chunk size, index settings, candidate counts) needs evaluation against real queries. A golden set of question and expected-source pairs, run in CI, is how you keep improvements from silently breaking the citation contract.