Context
For APHP (AP-HP, Assistance Publique – Hôpitaux de Paris), while working as System Analyst at IBM France (2023 – 2024, Paris), I designed the E-Reply System: an AI-augmented medical monitoring platform for the automated analysis of patient communications. It integrates LLMs, RAG and vector search to improve the relevance of responses and to assist medical teams, in an environment with very demanding security and data-confidentiality requirements.
Problem
Medical teams receive patient communications that need timely and relevant responses. Reading and answering every message by hand does not scale, but an automated system in a hospital context has to be accurate enough to help clinicians and confidential enough to be trusted with patient data. Speed matters because a slow reply is a real cost for the person waiting for it, and relevance matters because a wrong answer costs more than a late one.
Architecture decisions
LLMs and RAG to improve response relevance
LLMs are combined with retrieval-augmented generation, so that responses draw on retrieved material instead of on the model alone. The aim is to improve the relevance of responses and to assist the medical teams, not to replace them.
Trade-off: RAG adds a retrieval component to build and evaluate. In return, answers are grounded in retrieved content, which is easier to check than free generation.
Vespa AI vector search
I implemented Vespa AI vector search as the search layer behind the assistant, which reduced email response time by 70%.
Trade-off: a dedicated vector search engine is one more system to operate. It is chosen for retrieval quality and speed over a general-purpose store.
Security and confidentiality as design constraints
The platform was built in a highly demanding security and data-confidentiality environment, and those requirements shaped the design from the start rather than being checked at the end.
Trade-off: strict constraints narrow the choice of tools and hosting and slow experimentation. They also make the system acceptable to the people who are responsible for the data.
Outcome
Vespa AI vector search reduced email response time by 70%, and the platform assists medical teams with the automated analysis of patient communications.
Stack
- LLMs
- RAG
- Vespa AI
- Vector search