Services
Three ways I get enterprise AI into production
Pick the problem you have. Each area comes with concrete deliverables, and the design work happens with your engineers, not in a slide deck.
01
Enterprise LLM platform
One governed path from your products to foundation models, with secure access, retrieval that cites its sources and tenant isolation designed in from day one.
Read more: Enterprise LLM platform →
Typical deliverables
- Architecture Design Document and ADRs for the platform
- AI gateway design: model access, EU-only inference, guardrails and model routing
- RAG architecture: hybrid search, reranking, citation policy and multi-tenant isolation
- Evaluation and observability approach: golden-set gates, traces and prompt management
- Architecture reviews and compliance gates for the teams building on the platform
Text description of the diagram
Four product capabilities (a payroll declaration correction assistant, timesheet OCR, in-product assistants and a document-chasing agent) send all model requests to a central AI gateway. The gateway applies EU-only inference profiles, Guardrails PII filtering and model routing, then calls Amazon Bedrock. A service control policy enforces EU-only inference organisation-wide.
Text description of the diagram
A question reaches the platform retrieval API. Retrieval runs on Aurora PostgreSQL with pgvector, using an HNSW index and hybrid search with reranking, while Row-Level Security isolates each tenant inside the database. The model is called through the AI gateway, and a strict citation policy means every answer cites its sources or declines to answer.
02
Agents you can trust with real actions
Agents that act on your systems without ever exceeding the user’s permissions, with validation gates, an audit trail and regulatory alignment.
Read more: Agents you can trust with real actions →
Typical deliverables
- MCP action-server design: versioned tool manifests, human validation on writes, idempotency
- Identity design: permission inheritance through OAuth2 Token Exchange and workload identities
- Decision log and audit-export design so every AI decision can be replayed
- GDPR and AI Act by design: data residency, minimisation, erasure and transparency
- LLM evaluation gates in CI
Text description of the diagram
An agent calls the single MCP action server, which exposes versioned tool manifests. Each action passes precondition and idempotency checks and, for every write, human validation. Only then is the action executed. An append-only decision log makes every AI decision replayable, with a WORM export to S3 Object Lock. Throughout, the agent inherits the requesting user permissions through OAuth2 Token Exchange and never exceeds them.
03
Cloud and platform engineering
Highly available platforms on AWS and GCP, and the delivery pipeline to run them.
Read more: Cloud and platform engineering →
Typical deliverables
- Multi-AZ architecture and infrastructure as code (Terraform)
- Kubernetes, GitOps and CI/CD standardisation, up to a company-wide software factory
- Observability and quality gates
- Legacy-to-microservices migration guidance
- Architecture documentation (DAT, ADRs)
Text description of the diagram
Users reach the platform through load balancing across availability zones. Inside one AWS region, application instances run in more than one availability zone, which is what a multi-AZ design provides: the platform keeps serving when a single zone is unavailable. The infrastructure is provisioned as code (IaC), and the architecture ensured 99.9% availability.
How it works
How an engagement runs
Four steps, each with something you can hold in your hands.
-
Talk
A 30-minute call. You describe what you are building and what is blocking it. I tell you honestly whether I can help.
-
Map
I review what exists and design the target: a reference architecture and decision records that your team can challenge.
-
Build
I work inside your team on the hard parts: the gateway, retrieval, agent actions, the pipelines. You get code, not slides.
-
Hand over
Evaluation gates, runbooks and dashboards, so your team can run the platform and change it without me.
Frequently asked questions
What does an engagement usually produce?
Architecture documents (an Architecture Design Document and ADRs), reference designs, review gates, and hands-on technical leadership of the teams building the platform.
Which cloud and AI stacks do you work with?
AWS (Amazon Bedrock, EKS, Aurora PostgreSQL with pgvector, IAM and SCPs), GCP (Cloud Run, Pub/Sub), Kubernetes and Terraform, agent runtimes such as Mastra and LangGraph, MCP, DeepEval and Langfuse.
Can you work in regulated environments?
Yes. I have worked in HDS-compliant environments (Inserm) and on a medical follow-up platform with strict confidentiality requirements (APHP), and I design for GDPR and the AI Act by default.
How do you approach RAG?
Retrieval sits behind an owned API with a strict citation policy: every answer cites its sources or declines, and tenant isolation is enforced in the database with Row-Level Security.
Which languages do you work in?
French (native), English and Arabic (fluent).
How do we start?
Send me a message or book a meeting from the contact page. We talk through the problem, and I tell you what I would do first.
Last updated:
Got an AI project stuck between demo and production?
Tell me what you are building. After a first conversation you will know whether I can help and what I would do first.
+33 6 02 73 49 22 me@osmanrami.fr