Most PDPL conversations start in the legal department and arrive at the engineering team as a policy document. By then the expensive decisions have already been made — where the database lives, what the staging environment contains, how long logs are kept — and unpicking them costs more than designing for them would have.
Saudi Arabia’s Personal Data Protection Law came into force in September 2023 with a one-year grace period, and has been fully enforceable since 14 September 2024, supervised by SDAIA. If your software holds personal data about people in the Kingdom, the obligations are already live.
What follows is the engineering view: the decisions that are cheap before the build and painful after.
Residency is an architecture decision, not a hosting invoice
The provision with the most structural consequence is data localisation. Personal data of people in Saudi Arabia is expected to be processed inside the Kingdom, with transfers outside it permitted only under defined conditions and after an assessment of the risk.
Teams usually read that as “pick a Saudi region” and consider it handled. The region is the easy part. The parts that quietly leave the Kingdom are everywhere else in the stack:
- Managed logging and error-tracking services that default to a US or EU ingest endpoint.
- Analytics and session-replay tools that capture form fields.
- Email, SMS and push providers, which by definition carry personal data.
- Backups and disaster-recovery replicas, which are often configured to a different region on purpose.
- A third-party model API called from your application with customer text in the prompt.
Each one is a transfer. None of them appear in a diagram that only shows the primary database. The useful exercise is to list every destination a byte of personal data can reach, including the ones your vendors introduce, and then decide which are defensible.
Non-production environments are still processing
The most common gap we see has nothing to do with the production system. It is the copy.
A production dump restored into staging so a bug can be reproduced. An export sent to a developer’s laptop. A test suite seeded from last month’s real customers because synthetic data was too much work. Those are all processing of personal data, and the obligations do not switch off because the environment is labelled “dev”.
The fix is not a policy telling people not to do it, because under deadline pressure they will. The fix is making the compliant path the easy one: a seeded, synthetic dataset that is good enough to reproduce real bugs, and a masked-restore script that is faster to run than a raw copy. If the safe route is slower than the unsafe route, the unsafe route wins.
Deletion has to be designed in
Data-subject rights are where retrofitting hurts most. A request to delete someone’s data is simple to honour when the system was built expecting it, and close to impossible when it was not.
The hard cases are predictable: the row is referenced by a dozen foreign keys; the name is denormalised into an invoice PDF; the email address is in an event stream that was designed to be immutable; a copy sits in the analytics warehouse; the nightly backup from three months ago still contains everything.
Decide early what deletion means in your system — hard delete, crypto-shredding of a per-subject key, or anonymisation that genuinely cannot be reversed — and make the choice per data store rather than once for the whole platform. An append-only ledger and a user profile table need different answers.
Keep a record that survives being asked for
Accountability obligations assume you can show what you do, not merely assert it. In practice that means a record of processing activities that is accurate, and logs that establish who accessed what.
A record of processing that was written once during a compliance project and never updated is worse than useless, because it is now a document that contradicts your system. Tie it to something that changes when the system changes — a file in the repository reviewed at release, an entry required when a new third-party service is added — so drift is visible.
Consent is a data model, not a checkbox
Treating consent as a boolean on the user record fails the first time someone asks what exactly a customer agreed to in March.
What you need is a record per purpose, with the version of the notice shown, the timestamp, and the channel it was collected through — and the ability to withdraw one purpose without withdrawing the others. Marketing consent and service-communication consent are different things, and a single flag cannot represent a customer who wants invoices but not campaigns.
Where this lands on cost
None of this is exotic engineering. Residency-aware infrastructure, synthetic test data, a deletion strategy, a consent model and honest audit logging are all ordinary work. They are expensive only when they arrive late, because each one touches the data model, and the data model is what everything else is built on.
The practical sequence is to settle residency and the deletion strategy before the schema is fixed, and treat the rest as normal backlog.
We build systems that have to satisfy this from the first commit — see custom software development and managed cloud services, or our team in Madinah, Saudi Arabia. If you are building for the Kingdom, what “Arabic-first” actually means covers the other half of the problem.
Common questions
Does PDPL apply to a company based outside Saudi Arabia?
It can. The law reaches the processing of personal data of individuals inside the Kingdom, which means a foreign SaaS platform or service provider with Saudi users can fall within scope even with no local entity. Where you are incorporated matters less than whose data you hold.
Can we store Saudi personal data outside the Kingdom?
Not freely. Localisation is the default expectation, and transfers abroad are permitted only under defined conditions and after assessing the risk. Treat every destination — including logging, analytics, email and model APIs — as a transfer that needs a justification, not just your primary database region.
Do the rules apply to our staging and test environments?
Yes. Copying real customer data into a test system is still processing it. The practical answer is a synthetic dataset and a masked-restore path that developers find faster than taking a raw production dump, so the compliant route is also the convenient one.


