Context

ENSO’s GenAI programme runs on a shared, multi-tenant AI platform. It powers four product capabilities: a payroll declaration correction assistant, timesheet OCR, in-product assistants and an autonomous document-chasing agent. I am the lead architect. I authored the Architecture Design Document and I drive its enforcement across teams through architecture reviews, ADRs and compliance gates.

Problem

A shared AI platform has to settle the same questions once, for every product that runs on it: how models are accessed, how answers are grounded in sources, what an AI agent is allowed to do, how quality is kept from regressing, and how regulation is met. The architecture below answers each of them at platform level.

Architecture decisions

A single gateway to the models

All model access goes through a central AI gateway, the single access path to Amazon Bedrock. It enforces EU-only inference profiles organisation-wide through a service control policy (SCP), applies Bedrock Guardrails for PII filtering, and routes requests to models so that products are isolated from model lifecycle changes.

Trade-off: the gateway is one more component to operate and a single path to keep highly available. In exchange, data residency, filtering and model changes are governed in one place instead of in every product.

Retrieval that cites its sources or declines

RAG runs on Aurora PostgreSQL with pgvector, using HNSW indexes and hybrid search with reranking, behind a retrieval API the platform owns. A strict citation policy applies: every answer cites its sources or declines to answer. Tenant isolation is enforced inside the database with Row-Level Security, not in application code.

Trade-off: owning the retrieval API and pushing isolation into PostgreSQL ties the design to one database engine. It makes the guarantees testable at the data layer and gives every product a single retrieval contract.

Agentic actions behind one MCP action server

Agents act on business systems only through a single MCP (Model Context Protocol) action server. Tool manifests are versioned, every write requires human validation, and actions carry idempotency and precondition checks. An append-only decision log makes every AI decision replayable, with a WORM export to S3 Object Lock.

Trade-off: human validation on every write slows fully autonomous flows by design. It is a deliberate choice: writes stay reviewable, and the decision log lets any outcome be replayed and audited.

An identity design with no privilege escalation

The AI inherits the requesting user’s permissions through OAuth2 Token Exchange (Keycloak, OIDC), and workloads use IRSA identities on Kubernetes. The AI can do what the user can do, never more.

Trade-off: because permissions follow the user, an agent’s reach is only as broad as that user’s rights. That removes a whole class of privilege-escalation risk at the cost of some convenience.

A quality and observability loop

DeepEval golden-set gates block regressions in CI. Langfuse is self-hosted in-VPC on EKS for traces, prompts and online feedback. Delivery is GitOps with feature-flagged dark releases.

Trade-off: golden sets need curation and self-hosting Langfuse means operating it. The gain is that prompt and model changes ship through the same gated pipeline as code, and trace data stays inside the VPC.

One protocol boundary between TypeScript and C#

The platform is a hybrid TypeScript and C# architecture: the Mastra agent runtime and the official MCP SDKs, with the language seam placed on the MCP protocol boundary. Contract tests and a shared schema repository guard that seam.

Trade-off: two languages add tooling and skills overhead. Putting the seam on a protocol boundary keeps each side independently deployable and testable, and the contract tests catch drift early.

GDPR and the AI Act by design

The design keeps data in the EU, minimises the data it handles, coordinates the right to erasure across vectors, traces and audit records, and makes AI use transparent to end users.

Trade-off: erasure has to reach every store that holds derived data, which is more work than deleting source rows. Designing it in from the start is cheaper than retrofitting it.

Outcome

Four product capabilities run on one shared, governed platform. The architecture is set out in the Architecture Design Document, and architecture reviews, ADRs and compliance gates enforce it across teams.

Stack

  • Amazon Bedrock
  • Aurora PostgreSQL
  • pgvector
  • EKS
  • Keycloak
  • MCP
  • Mastra
  • TypeScript
  • C#
  • DeepEval
  • Langfuse