An enterprise AI platform is the shared foundation that lets several products use language models without each one re-solving security, grounding, control and quality. It has five layers: an AI gateway for model access, retrieval that cites its sources, a controlled action layer for agents, evaluation and observability, and identity with tenancy and compliance. This guide describes each layer, the order to build them in and the mistakes that cost the most.

Why a platform instead of one integration per product

The first AI feature in a company is usually a direct integration with a model provider, and it works. The second and third features copy it. By the fifth, there are five ways of choosing a model, five ways of handling personal data, five answers to “where do the answers come from?” and no place to change any of them.

A platform turns those decisions from per-product choices into platform choices. Products depend on a small, stable contract. The organisation gets one place to enforce policy, observe usage and improve quality, and a model change stops being a project.

Layer 1: the AI gateway

Every call to a foundation model goes through one service. The gateway enforces data residency (for example EU-only inference profiles on Amazon Bedrock), filters personal data with guardrails and routes each capability to a concrete model, so that products ask for “summarise” or “extract” rather than naming a model. An account-level control policy covers traffic that bypasses the gateway. The details are in the AI gateway guide.

Layer 2: retrieval that cites its sources

Retrieval-augmented generation only earns trust if answers carry citations and the system declines when it has none. Keep retrieval behind an API the platform owns, run hybrid search with reranking, store embeddings next to the relational data (PostgreSQL with pgvector is a sound default) and enforce tenant isolation in the database with Row-Level Security. The contract and its limits are covered in the RAG guide.

Layer 3: agent actions under control

Reading is low risk and writing is not. Route every agent action through a single MCP action server with versioned tool manifests, human approval on each write, idempotency and preconditions, and an append-only decision log. The agent holds no rights of its own: it acts with the requesting user’s permissions. See the agent actions guide.

Layer 4: evaluation and observability

A language-model application changes whenever a prompt, a model or a retrieval setting changes. Quality therefore has to be measured on every change: golden sets, evaluation gates in CI and traces of production calls that feed back into the sets. The method is described in the evaluation guide.

Layer 5: identity, tenancy and compliance

Identity ties the layers together. Users authenticate once, the platform exchanges their token for scoped tokens downstream, workloads use short-lived cloud identities, and tenants are isolated in the data layer. Compliance is designed in, not audited afterwards: data residency, minimisation, erasure and transparency for GDPR, and the obligations of the AI Act that apply to the use case.

The order to build it in

Start with the gateway: it is small, it removes the most duplicated work and it gives you usage data. Add retrieval next, because most first use cases need grounded answers. Put evaluation in place before the third product ships, not after the first incident. Introduce agent actions last, and only once identity and the decision log exist, because a write-capable agent without them is the riskiest thing you can deploy.

Failure modes to avoid

  • A platform nobody asked for. Build the layer a real product needs first and generalise from it.
  • A gateway that becomes a bottleneck. Keep its contract small and stable, and make it highly available.
  • Policy only in prompts. Enforce rules in code and infrastructure (citation checks, database policies, account-level controls), because prompts can be talked around.
  • No evaluation until something breaks. Without a golden set, every improvement is an opinion.
  • Agents with their own privileges. An agent should never be able to do what the user cannot.

In practice

On ENSO’s shared, multi-tenant platform I designed these layers together: a central AI gateway on Amazon Bedrock, RAG on Aurora PostgreSQL with pgvector behind an owned retrieval API, agent actions through one MCP action server with human validation, evaluation gates in CI and identity through Keycloak and OAuth2 Token Exchange. The platform powers four product capabilities. The full case is in the ENSO case study, and the services behind it are described on the enterprise LLM platform page.

Trade-offs and limits

A platform is an investment that pays off with the second and third product, not the first. It adds services to run and a team or role to own them. A single product with a single model may not need it yet. Centralising also concentrates risk: the gateway and the retrieval API must be reliable, and they need an owner who keeps up with the products that depend on them.