Data privacy
Your data does not leave without you knowing where it goes.
It is the first objection directors put to me whenever generative artificial intelligence comes up, and it is a fair one. There is no single answer to it: there are four, depending on how sensitive your material is.
Where to start
The six questions to put to any supplier.
Including me. A vague answer to any one of them is reason enough to look elsewhere.
- 01
Where is the data processed?
You check the hosting region in the contract, not on a marketing page. A company headquartered outside Europe may well process inside Europe, and the reverse is just as true.
- 02
Is it kept, and for how long?
Many services hold on to exchanges for a few days for security reasons. That is often acceptable, provided you know about it and can ask for zero retention where it matters.
- 03
Is it used to train a model?
This is the sharpest dividing line between consumer use and business use under contract. In a business setting the answer has to be no, and it has to be in writing.
- 04
Which subprocessors are involved?
A supplier who cannot tell you who, below them, sees your data cannot guarantee you much. The list should be handed over and kept up to date.
- 05
Who at the supplier can read the data?
There is almost always an access procedure for incidents. The point is not to abolish it but to know it exists, on what conditions it triggers, and whether it leaves a trace.
- 06
What happens if you stop?
Getting your data back in a usable format, erasure at the supplier, and ownership of the code produced. All three are settled at the start, never on the way out.
The scale
Four levels of sensitivity, four technical answers.
Most companies are in fact handling several levels at once. The work is to avoid applying the constraint of the highest level to everything, which comes down to doing nothing anywhere.
| Level | What you handle | Examples | The technical answer |
|---|---|---|---|
| Level 1 | Public data, or data with nothing at stake | Published content, documentation, catalogues, market data | A programming interface from a supplier under contract. No particular precaution beyond a sound contract. |
| Level 2 | Company confidential, no personal data | Proposals, supplier contracts, pricing, internal methods | An interface under contract, with no training on your data, retention configured, and European hosting verified. |
| Level 3 | Personal data about customers or staff | Customer files, call recordings, personnel records | The same framework, plus GDPR compliance: register, legal basis, retention periods, informing the people concerned. Often, work on cut-down or anonymised data. |
| Level 4 | Regulated or strategic data | Health, defence, trade secrets, anything under a strict confidentiality agreement | An open model running on infrastructure you control, or an architecture that strictly separates what may leave from what must not. |
The self-assessment asks which level you are on and adapts what it recommends. Start it
In practice
How I work, whatever the level.
-
We separate what leaves from what stays
Most builds have no need to send the whole database to a model. We isolate the strict minimum, often a cut-down or anonymised extract, and the rest never leaves your systems.
-
You choose where it runs
On your premises, on your cloud, or on mine for a transitional period. The choice follows the sensitivity level and whether you have someone in house able to run the system.
-
The data note is part of the diagnostic
Within the ten days of the diagnostic, every lever identified comes with its answer: what data would leave, where it would be processed, and what that means in law.
-
The code and the data are yours
The code goes into a repository you own, from the first day. No dependency on a service that only I could administer.
To do at your end
The checklist, worth running even without me.
Eight points. They take half a day, and they are worth more than a twenty-page policy nobody will read.
The full guide- 01
List the generative AI already in use across the business, including the uses nobody declared
- 02
Rate each use against a four-level sensitivity scale
- 03
Check, tool by tool, whether it trains a model on your data, and write the answer down
- 04
Check the hosting region and the retention period in the contract, not on the website
- 05
Enter the processing concerned in your register, with its legal basis and its retention period
- 06
Inform the people concerned whenever personal data is being processed
- 07
Write one plain rule, readable by everyone, on what may be given to a tool and what may not
- 08
Name someone accountable, without which the rule exists only on paper
Data privacy
The questions that keep coming back
Can I use generative AI with personal data?
Yes, provided you treat it like any other processing of personal data: a legal basis, an entry in the register, a retention period, information given to the people concerned, and a subprocessor under contract. What is not allowed is doing it with none of that in place.
Is a model running on our own machines the only genuinely safe option?
No, and it is often a bad trade. An open model on your own hardware costs equipment and skills, and it stays behind the best models on the market. That option earns its place at level four, or when a client contract requires it. Below that, an interface under contract, properly framed, is both safer in practice and more effective.
Are consumer tools a problem?
The problem is not the tool, it is the absence of a rule. Staff pasting internal documents into a personal service, with no contract and with nobody aware of it, create a real risk. The answer is not a ban, which never holds, but framed access plus a plain statement of what is allowed.
What does the European AI Act say?
For a company automating internal tasks, the substance comes down to two practical duties: know what AI use exists inside your business, and be transparent when someone is dealing with an automated system or reading generated content. The heavy obligations target high-risk uses, above all anything touching recruitment and the assessment of individuals, where you need legal advice.
How do I check what a supplier claims?
Ask for the documents rather than the statements: the processing agreement, the list of subprocessors, the hosting region, the retention period, and the policy on training with customer data. A serious supplier hands them over without fuss.
This page describes engineering practice, not legal advice. For a precise assessment of your own processing, in particular any use touching recruitment or the assessment of individuals, take proper professional advice. If you want to talk a specific case through, write to me at contact@praedic.com .
Next step
Worried about your data? That is the right question.
The diagnostic settles it lever by lever, before a single line of code is written. You leave with a written note on what would go out, and where it would go.
Diagnostic
Étape 110 days
€9,500 excl. VAT · fixed fee, one area of the business
- A map of where the time actually goes in the area audited
- Each lever costed in hours freed, ranked by effort and by gain
- The data note: what leaves your systems, and where it is processed
- A three-month plan, naming the first system to build