One AI platform, four products, one set of rules
Lead architect of the shared platform behind four AI product capabilities, from model access to agent actions.
- Amazon Bedrock
- Aurora PostgreSQL
- pgvector
- EKS
Services
An enterprise LLM platform gives every product one governed way to use foundation models: a central AI gateway for model access, retrieval that cites its sources, tenant isolation in the data layer and evaluation gates in CI. I design it end to end on AWS, and I am doing exactly that today as lead architect of ENSO’s shared AI platform.
Without a platform, each product team integrates models on its own. Every team then solves the same problems again: which regions and models are allowed, how personal data is filtered, how answers are grounded in sources, how quality is measured and what happens when a model is retired.
A platform settles these questions once. Products call a stable contract, and the platform decides where, how and with which model the request is served. A model change becomes a configuration change in one place instead of a change in every product.
Four product capabilities (a payroll declaration correction assistant, timesheet OCR, in-product assistants and a document-chasing agent) send all model requests to a central AI gateway. The gateway applies EU-only inference profiles, Guardrails PII filtering and model routing, then calls Amazon Bedrock. A service control policy enforces EU-only inference organisation-wide.
A question reaches the platform retrieval API. Retrieval runs on Aurora PostgreSQL with pgvector, using an HNSW index and hybrid search with reranking, while Row-Level Security isolates each tenant inside the database. The model is called through the AI gateway, and a strict citation policy means every answer cites its sources or declines to answer.
One path to Amazon Bedrock with EU-only inference profiles, PII filtering through Guardrails and model routing, backed by an organisation-wide service control policy.
RAG on PostgreSQL with pgvector, hybrid search and reranking behind an API the platform owns, with a citation policy: every answer cites its sources or declines.
Row-Level Security policies so that each tenant only queries its own rows, tested at the data layer.
Golden sets and evaluation gates in CI, with traces and prompt management, so a regression fails the build instead of reaching users.
An Architecture Design Document and decision records, enforced through architecture reviews and compliance gates, with GDPR and AI Act concerns handled in the design.
How it works
Four steps, each with something you can hold in your hands.
A 30-minute call. You describe what you are building and what is blocking it. I tell you honestly whether I can help.
I review what exists and design the target: a reference architecture and decision records that your team can challenge.
I work inside your team on the hard parts: the gateway, retrieval, agent actions, the pipelines. You get code, not slides.
Evaluation gates, runbooks and dashboards, so your team can run the platform and change it without me.
Lead architect of the shared platform behind four AI product capabilities, from model access to agent actions.
Vespa AI vector search cut email response time by 70% on a medical follow-up platform built with LLMs and RAG.
A reference architecture for an enterprise AI platform: AI gateway, retrieval, agent actions, evaluation and identity, and the order to build them in.
A design guide to an enterprise AI gateway on Amazon Bedrock: EU-only inference profiles, SCP enforcement, Guardrails PII filtering and model routing.
How to build RAG that cites its sources or declines to answer, on PostgreSQL with pgvector, HNSW, hybrid search, reranking and Row-Level Security.
How to evaluate LLM applications in CI: build golden sets, choose metrics, gate releases with DeepEval and trace production with Langfuse.
An AI gateway is the single service through which every product reaches foundation models. It enforces policy in one place (allowed regions, personal-data filtering, model routing), records usage and lets models change without changing products.
So that the citation policy, tenant isolation and search quality are enforced once. Products get one retrieval contract, and every improvement to search benefits all of them.
Mostly AWS: Amazon Bedrock, EKS, Aurora PostgreSQL with pgvector, IAM and service control policies, plus Keycloak for identity. I also have production experience on GCP.
Yes. I design for EU-only inference, data minimisation and auditability from the start, and I have worked in HDS-compliant and medical settings.
Last updated:
Tell me what your teams are building. After one call you will know whether I can help and what I would review first.
+33 6 02 73 49 22 me@osmanrami.fr