Where does your data go? Data residency for AI systems in the UAE
Operators ask us this before anything else, and they are right to. Here is what managed cloud, your cloud and on-premise really mean when the system includes a language model, and how to decide.
Armaan PuriFounder and CEO, CrafterTech AI Middle East8 October 2026 · 7 min read
Every serious conversation we have with a UAE enterprise reaches the same question within the first ten minutes: where does our data go? It is the right question. An AI system that reads your tenders, your manuals or your customer conversations is handling the most sensitive material in the business, and a vague answer should end the meeting.
The honest answer has three parts: where the data is stored, where it is processed, and which model sees it. Most vendors answer the first and go quiet on the other two.
The rules that apply
The UAE's federal Personal Data Protection Law sets the baseline for personal data, with its own rules on cross-border transfer. The financial free zones, DIFC and ADGM, run their own data protection regimes. Government and semi-government entities, including the national energy companies, usually add contractual requirements of their own, and some classes of data are expected to stay inside the country full stop. Sector regulators add more. None of this is exotic; it simply has to be designed in, not promised later.
Three deployment shapes, plainly
- Managed cloud. We host and run the system for you, in a cloud region you approve, under contractual controls. Fastest to start. Right for workflows that do not touch restricted data.
- Your cloud. The whole system runs inside your own AWS, Azure or Google Cloud account, in the UAE regions those providers operate, under your keys and your identity system. We deploy and maintain; you own the environment.
- On-premise. The system, including the language model, runs on servers inside your network. Nothing leaves. Slower to stand up, and the only correct answer for some data.
The part people miss: the model
Storage location is the easy part. The harder question is which model processes the text and where that model runs. A commercial model called through an API processes your data on the provider's infrastructure, in whatever region that provider offers. An open-source model can run inside your own cloud account or on your own hardware, so the data never leaves your boundary.
This is why we choose models per task rather than per vendor. A tender desk might use a commercial model for general reasoning on documents that are not restricted, and an open model hosted inside your environment for anything that is. The deployment decision is made clause by clause, and it is written down.
Questions to ask any vendor
- In which country and region is our data stored, and who holds the keys?
- Which model processes our documents, and where does that model run?
- Is any of our data used to train or improve a model? The answer must be no, in writing.
- What is logged, who can read the logs, and how long are they kept?
- Can we leave? How is our data returned and deleted at the end?
If a vendor cannot answer these in a page, they have not designed for them.
How we decide with you
Every deployment starts with a diagnostic of one workflow. Part of that diagnostic is a data map: what the system will read, how sensitive each source is, and which of the three shapes each source requires. Most organisations end up with a mix, and that is fine. The system is built to run the same way in all three; only the boundary moves.
Controls are the same regardless of where the system runs. Role-based access decides what each person can see, every consequential action waits for approval, and an audit trail records what was read and decided. Residency keeps the data in the right place. Governance keeps the right people in charge of it. You need both.
Map one operational workflow.
Tell us where work is slow, manual or risky. We map it end to end and show you what a deployed system would change.