Services

Three ways I get enterprise AI into production

Pick the problem you have. Each area comes with concrete deliverables, and the design work happens with your engineers, not in a slide deck.

01

Enterprise LLM platform

One governed path from your products to foundation models, with secure access, retrieval that cites its sources and tenant isolation designed in from day one.

Read more: Enterprise LLM platform →

Typical deliverables

  • Architecture Design Document and ADRs for the platform
  • AI gateway design: model access, EU-only inference, guardrails and model routing
  • RAG architecture: hybrid search, reranking, citation policy and multi-tenant isolation
  • Evaluation and observability approach: golden-set gates, traces and prompt management
  • Architecture reviews and compliance gates for the teams building on the platform
AI gateway architecture Product capabilities Payroll correction Timesheet OCR In-product assistants Document-chasing agent AI gateway EU-only inference profilesGuardrails PII filteringModel routing Amazon Bedrock EU inference SCP: EU-only inference, organisation-wide
Every product reaches Amazon Bedrock through one gateway that enforces EU-only inference, PII guardrails and model routing.
Text description of the diagram

Four product capabilities (a payroll declaration correction assistant, timesheet OCR, in-product assistants and a document-chasing agent) send all model requests to a central AI gateway. The gateway applies EU-only inference profiles, Guardrails PII filtering and model routing, then calls Amazon Bedrock. A service control policy enforces EU-only inference organisation-wide.

RAG with citations architecture Question Retrieval API Owned by the platform Aurora PostgreSQL + pgvector HNSW indexHybrid search + reranking Row-Level Security isolates tenants LLM via the AI gateway Every answer cites its sourcesor declines to answer
Retrieval sits behind an owned API, runs on PostgreSQL with pgvector and isolates tenants in the database. Every answer cites its sources or declines.
Text description of the diagram

A question reaches the platform retrieval API. Retrieval runs on Aurora PostgreSQL with pgvector, using an HNSW index and hybrid search with reranking, while Row-Level Security isolates each tenant inside the database. The model is called through the AI gateway, and a strict citation policy means every answer cites its sources or declines to answer.

02

Agents you can trust with real actions

Agents that act on your systems without ever exceeding the user’s permissions, with validation gates, an audit trail and regulatory alignment.

Read more: Agents you can trust with real actions →

Typical deliverables

  • MCP action-server design: versioned tool manifests, human validation on writes, idempotency
  • Identity design: permission inheritance through OAuth2 Token Exchange and workload identities
  • Decision log and audit-export design so every AI decision can be replayed
  • GDPR and AI Act by design: data residency, minimisation, erasure and transparency
  • LLM evaluation gates in CI
Human-in-the-loop agent action flow Agent MCP action server Versioned tool manifests Precondition + idempotency checks Human validation on every write Action executed Append-only decision log WORM export to S3 Object Lock Identity: OAuth2 Token Exchange (Keycloak)Inherits the user’s permissions, never more
An agent can only act through one MCP action server. Every write passes checks and human validation, and every decision is logged for replay.
Text description of the diagram

An agent calls the single MCP action server, which exposes versioned tool manifests. Each action passes precondition and idempotency checks and, for every write, human validation. Only then is the action executed. An append-only decision log makes every AI decision replayable, with a WORM export to S3 Object Lock. Throughout, the agent inherits the requesting user permissions through OAuth2 Token Exchange and never exceeds them.

03

Cloud and platform engineering

Highly available platforms on AWS and GCP, and the delivery pipeline to run them.

Read more: Cloud and platform engineering →

Typical deliverables

  • Multi-AZ architecture and infrastructure as code (Terraform)
  • Kubernetes, GitOps and CI/CD standardisation, up to a company-wide software factory
  • Observability and quality gates
  • Legacy-to-microservices migration guidance
  • Architecture documentation (DAT, ADRs)
Multi-AZ AWS platform layout Users Load balancing across zones AWS region Availability zone A Availability zone B Application Application Provisioned as code (IaC) Ensuring 99.9% availability
A web platform spread across AWS availability zones and provisioned as code, ensuring 99.9% availability.
Text description of the diagram

Users reach the platform through load balancing across availability zones. Inside one AWS region, application instances run in more than one availability zone, which is what a multi-AZ design provides: the platform keeps serving when a single zone is unavailable. The infrastructure is provisioned as code (IaC), and the architecture ensured 99.9% availability.

How it works

How an engagement runs

Four steps, each with something you can hold in your hands.

  1. Talk

    A 30-minute call. You describe what you are building and what is blocking it. I tell you honestly whether I can help.

  2. Map

    I review what exists and design the target: a reference architecture and decision records that your team can challenge.

  3. Build

    I work inside your team on the hard parts: the gateway, retrieval, agent actions, the pipelines. You get code, not slides.

  4. Hand over

    Evaluation gates, runbooks and dashboards, so your team can run the platform and change it without me.

Frequently asked questions

What does an engagement usually produce?

Architecture documents (an Architecture Design Document and ADRs), reference designs, review gates, and hands-on technical leadership of the teams building the platform.

Which cloud and AI stacks do you work with?

AWS (Amazon Bedrock, EKS, Aurora PostgreSQL with pgvector, IAM and SCPs), GCP (Cloud Run, Pub/Sub), Kubernetes and Terraform, agent runtimes such as Mastra and LangGraph, MCP, DeepEval and Langfuse.

Can you work in regulated environments?

Yes. I have worked in HDS-compliant environments (Inserm) and on a medical follow-up platform with strict confidentiality requirements (APHP), and I design for GDPR and the AI Act by default.

How do you approach RAG?

Retrieval sits behind an owned API with a strict citation policy: every answer cites its sources or declines, and tenant isolation is enforced in the database with Row-Level Security.

Which languages do you work in?

French (native), English and Arabic (fluent).

How do we start?

Send me a message or book a meeting from the contact page. We talk through the problem, and I tell you what I would do first.

Last updated:

Got an AI project stuck between demo and production?

Tell me what you are building. After a first conversation you will know whether I can help and what I would do first.

+33 6 02 73 49 22 me@osmanrami.fr