Services

Enterprise LLM platform architecture

An enterprise LLM platform gives every product one governed way to use foundation models: a central AI gateway for model access, retrieval that cites its sources, tenant isolation in the data layer and evaluation gates in CI. I design it end to end on AWS, and I am doing exactly that today as lead architect of ENSO’s shared AI platform.

What an enterprise LLM platform is

Without a platform, each product team integrates models on its own. Every team then solves the same problems again: which regions and models are allowed, how personal data is filtered, how answers are grounded in sources, how quality is measured and what happens when a model is retired.

A platform settles these questions once. Products call a stable contract, and the platform decides where, how and with which model the request is served. A model change becomes a configuration change in one place instead of a change in every product.

AI gateway architecture Product capabilities Payroll correction Timesheet OCR In-product assistants Document-chasing agent AI gateway EU-only inference profilesGuardrails PII filteringModel routing Amazon Bedrock EU inference SCP: EU-only inference, organisation-wide
Every product reaches Amazon Bedrock through one gateway that enforces EU-only inference, PII guardrails and model routing.
Text description of the diagram

Four product capabilities (a payroll declaration correction assistant, timesheet OCR, in-product assistants and a document-chasing agent) send all model requests to a central AI gateway. The gateway applies EU-only inference profiles, Guardrails PII filtering and model routing, then calls Amazon Bedrock. A service control policy enforces EU-only inference organisation-wide.

RAG with citations architecture Question Retrieval API Owned by the platform Aurora PostgreSQL + pgvector HNSW indexHybrid search + reranking Row-Level Security isolates tenants LLM via the AI gateway Every answer cites its sourcesor declines to answer
Retrieval sits behind an owned API, runs on PostgreSQL with pgvector and isolates tenants in the database. Every answer cites its sources or declines.
Text description of the diagram

A question reaches the platform retrieval API. Retrieval runs on Aurora PostgreSQL with pgvector, using an HNSW index and hybrid search with reranking, while Row-Level Security isolates each tenant inside the database. The model is called through the AI gateway, and a strict citation policy means every answer cites its sources or declines to answer.

What gets designed

  1. AI gateway

    One path to Amazon Bedrock with EU-only inference profiles, PII filtering through Guardrails and model routing, backed by an organisation-wide service control policy.

  2. Retrieval

    RAG on PostgreSQL with pgvector, hybrid search and reranking behind an API the platform owns, with a citation policy: every answer cites its sources or declines.

  3. Tenant isolation

    Row-Level Security policies so that each tenant only queries its own rows, tested at the data layer.

  4. Evaluation and observability

    Golden sets and evaluation gates in CI, with traces and prompt management, so a regression fails the build instead of reaching users.

  5. Governance

    An Architecture Design Document and decision records, enforced through architecture reviews and compliance gates, with GDPR and AI Act concerns handled in the design.

Typical deliverables

  • Architecture Design Document and ADRs for the platform
  • AI gateway design: model access, EU-only inference, guardrails and model routing
  • RAG architecture: hybrid search, reranking, citation policy and multi-tenant isolation
  • Evaluation and observability approach: golden-set gates, traces and prompt management
  • Architecture reviews and compliance gates for the teams building on the platform

Who this is for

  • Engineering and product leaders whose teams are building several AI features and need a shared foundation that can be reviewed.
  • Teams working under data-residency, GDPR or AI Act constraints who have to show how AI behaviour is controlled.
  • Organisations moving from a convincing pilot to a platform that has to pass security review and carry real users.

How it works

How an engagement runs

Four steps, each with something you can hold in your hands.

  1. Talk

    A 30-minute call. You describe what you are building and what is blocking it. I tell you honestly whether I can help.

  2. Map

    I review what exists and design the target: a reference architecture and decision records that your team can challenge.

  3. Build

    I work inside your team on the hard parts: the gateway, retrieval, agent actions, the pipelines. You get code, not slides.

  4. Hand over

    Evaluation gates, runbooks and dashboards, so your team can run the platform and change it without me.

Case studies

ENSO

2026 – Present

One AI platform, four products, one set of rules

Lead architect of the shared platform behind four AI product capabilities, from model access to agent actions.

  • Amazon Bedrock
  • Aurora PostgreSQL
  • pgvector
  • EKS
APHP

2023 – 2024

Cutting patient email response time by 70%

Vespa AI vector search cut email response time by 70% on a medical follow-up platform built with LLMs and RAG.

  • LLMs
  • RAG
  • Vespa AI
  • Vector search

Guides

Common questions

What is an AI gateway?

An AI gateway is the single service through which every product reaches foundation models. It enforces policy in one place (allowed regions, personal-data filtering, model routing), records usage and lets models change without changing products.

Why put retrieval behind an API the platform owns?

So that the citation policy, tenant isolation and search quality are enforced once. Products get one retrieval contract, and every improvement to search benefits all of them.

Which stack do you build it on?

Mostly AWS: Amazon Bedrock, EKS, Aurora PostgreSQL with pgvector, IAM and service control policies, plus Keycloak for identity. I also have production experience on GCP.

Can it run in a regulated environment?

Yes. I design for EU-only inference, data minimisation and auditability from the start, and I have worked in HDS-compliant and medical settings.

Last updated:

Designing an enterprise LLM platform?

Tell me what your teams are building. After one call you will know whether I can help and what I would review first.

+33 6 02 73 49 22 me@osmanrami.fr