AI architect in Paris · Enterprise AI & cloud

Take your AI from demo to production.

I design and build the platforms behind enterprise AI: model gateways, RAG that cites its sources and agents that act within your rules. Hands-on, on AWS, for teams that have to pass security review.

One request through the platform (illustrative)

  1. request an assistant asks a question
  2. guardrails personal data masked
  3. retrieval tenant-scoped sources found
  4. model routed inside the EU
  5. answer cites its sources

Delivered for and with

What I fix

Where enterprise AI projects get stuck

The model is rarely what stalls an AI project. It is the things around it: access, trust, control and the cloud underneath. Those are the parts I build.

Every team wires up its own model.

One governed door to every model

A central AI gateway on Amazon Bedrock enforces EU-only inference and PII filtering, then routes each request to the right model. Products stop depending on any single one.

  • Amazon Bedrock
  • Guardrails
  • SCP
  • Model routing
How it works

The assistant makes things up.

Answers that cite their sources, or decline

RAG on PostgreSQL with pgvector, hybrid search and reranking. Tenant isolation is enforced in the database itself.

  • pgvector
  • HNSW
  • Hybrid search
  • Row-Level Security
How it works

Nobody trusts an agent to act.

Agents that ask before they act

One MCP action server with versioned tools, human approval on every write and a decision log you can replay. The agent never holds more rights than the user.

  • MCP
  • OAuth2 Token Exchange
  • Keycloak
  • S3 Object Lock
How it works

Nobody can tell if the last change made it better.

Quality you can measure, in CI

Golden sets and evaluation gates in the pipeline, plus tracing of every prompt and answer, so a regression fails the build instead of reaching users.

  • DeepEval
  • Langfuse
  • Golden sets
How it works

It runs on one server. Then real traffic arrives.

Cloud foundations that stay up

Multi-AZ platforms on AWS and GCP defined as code, with Kubernetes, GitOps and CI/CD that your teams can run without me.

  • AWS
  • Kubernetes
  • Terraform
  • GitOps
How it works

You want to run open-source models yourself.

Your own GPUs, open-source models

Distributed inference on a GPU cluster, with the GenAI pipelines, fine-tuning and automation that go around it.

  • RunPod
  • ComfyUI
  • n8n
  • Fine-tuning
How it works

Toolbox

The stack I work in

Technologies from the platforms above and from my own projects.

  • Amazon Bedrock
  • AWS
  • Kubernetes
  • Terraform
  • PostgreSQL
  • pgvector
  • MCP
  • Keycloak
  • Google Cloud
  • Docker
  • Argo CD
  • Grafana
  • LangGraph
  • Mastra
  • DeepEval
  • Langfuse
  • Vespa
  • n8n
  • Spring Boot
  • TypeScript
  • Kafka
  • GitLab
  • Redis
  • Ansible

Selected work

Real platforms, real constraints

Six engagements, each told by the problem and the outcome.

Inserm

2025 – 2026

One software factory for 12 teams

40+ projects on automated CI/CD and 50% fewer manual interventions, on HDS-compliant infrastructure.

  • GitLab Enterprise
  • Nexus
  • SonarQube
  • Kubernetes
Volkswagen Group

2025 – 2026

99.9% availability, by design

Multi-AZ AWS infrastructure defined as code, ensuring 99.9% availability for the GENWIN web platform.

  • AWS
  • Multi-AZ
  • Infrastructure as code
  • Clean Architecture
SNCF

2025 – 2026

A data marketplace with 90,000 users

Spring Boot microservices on managed GCP services, with observability built in from the start.

  • Spring Boot
  • Microservices
  • Google Cloud Platform
  • Managed services
APHP

2023 – 2024

Cutting patient email response time by 70%

Vespa AI vector search cut email response time by 70% on a medical follow-up platform built with LLMs and RAG.

  • LLMs
  • RAG
  • Vespa AI
  • Vector search
Magik AI

2020 – 2025

Open-source AI on a GPU cluster

An interactive media generation platform running open-source models on a RunPod GPU cluster, with GenAI pipelines, agents and n8n workflows.

  • RunPod
  • Open-source models
  • RAG
  • Fine-tuning

All case studies →

How it works

How an engagement runs

Four steps, each with something you can hold in your hands.

  1. Talk

    A 30-minute call. You describe what you are building and what is blocking it. I tell you honestly whether I can help.

  2. Map

    I review what exists and design the target: a reference architecture and decision records that your team can challenge.

  3. Build

    I work inside your team on the hard parts: the gateway, retrieval, agent actions, the pipelines. You get code, not slides.

  4. Hand over

    Evaluation gates, runbooks and dashboards, so your team can run the platform and change it without me.

Field notes

What building these systems taught me

Design guides with the decisions and the trade-offs, not the hype.

All articles →

Last updated:

Got an AI project stuck between demo and production?

Tell me what you are building. After a first conversation you will know whether I can help and what I would do first.

+33 6 02 73 49 22 me@osmanrami.fr