Data privacy 6 min read

Using generative AI without exposing company data

The six questions to put to any supplier, a four-level sensitivity scale, what the GDPR actually requires, and a checklist to run before you go live.

This is the question that stops more projects than any other, and it is a good question. Knowing where your data goes when an employee opens a chat assistant deserves a precise answer rather than a sales promise. This guide gives you the questions to ask, a scale for sorting your data, and what the rules actually require.

The real risk is not the one people fear

The most widespread fear is of a model that learns your secrets and hands them to a competitor. That risk is real but has become marginal with serious suppliers, because professional offers rule it out explicitly by contract.

The everyday risk, the one that actually occurs inside companies, is far more prosaic. An employee, with no tool provided by the company, opens a personal account on a consumer service and pastes in a contract, a customer file or a payslip to get some help. There is no bad intent. There is simply nothing else available to them.

That observation dictates the strategy. Banning without providing an alternative protects nothing: it moves the usage out of your sight. Real protection comes from putting an approved tool in people’s hands, together with a rule simple enough to remember.

The difference between a consumer app and a programming interface

This distinction is the most useful thing in the guide, and it is poorly understood.

When you use the free or personal version of an online assistant, you are in a consumer setting. The terms of use may well provide that your exchanges help improve the service, and you have signed no data processing contract with the supplier. That is perfectly acceptable for drafting a text with nothing at stake. It is not acceptable for your customer data.

When a supplier builds a system for you, they do not use that app. They call a programming interface, meaning technical access to the model under a professional agreement. The terms there are different: content sent is not used to train the models, a processing agreement under the General Data Protection Regulation is signed, and the hosting region can often be chosen.

The same model, the same technology, two unrelated legal settings. So never judge a technology on the experience you have had of its consumer app.

There is a third setting, intermediate and very common: the professional offers sold by software publishers, bought as licences for your teams. They generally include a commitment not to reuse your data. They still have to be checked one by one, because terms vary between publishers and change over time.

The six questions to put to any supplier or publisher

Ask them in writing, and expect written answers. A serious supplier answers in a few lines without discomfort. A supplier who dodges them has told you just as much.

  • Where is my data processed, in which country and under which jurisdiction.
  • Is it retained after processing, and if so for exactly how long.
  • Is it used to train or improve the model, and is that commitment in the contract.
  • Which hosting region can I choose, and is that option included or charged for.
  • Who are your subprocessors, including the model supplier itself and the hosting provider.
  • What happens if I terminate: is my data deleted, within what period, and with what proof.

The fifth question is the one most often forgotten, and the most revealing. Many software publishers have added generative artificial intelligence features by calling a third-party model. Your contract is with the publisher, but your data passes through somebody else, sometimes outside Europe. That is not necessarily a problem, provided you know about it and the chain is documented.

A four-level sensitivity scale

Classifying all your data at the strictest level amounts to doing nothing. Telling the levels apart lets you move quickly where you can and carefully where you must.

Level 1, public or unremarkable data

Your published content, your sales material, a generic text, an internal note with nothing at stake. No particular constraint. Any serious professional tool will do, and depriving yourself of generative AI at this level makes no sense.

Level 2, internal non-personal data

Your procedures, your technical documents, your aggregated activity figures, your application code. Disclosure would be annoying without being grave. The sensible level here is a professional offer under contract, with a written commitment not to reuse the data and, preferably, hosting in Europe.

Level 3, personal data and customer data

Customer files, contracts, named correspondence, meeting transcripts, job applications. The General Data Protection Regulation applies in full. You need a proper processing agreement, European hosting, a short and defined retention period, an entry in your record of processing activities, and information given to the people concerned. This is the level of the great majority of worthwhile projects in a mid-sized company, and it is entirely workable.

Level 4, sensitive or strategic data

Health data, critical infrastructure data, industrial secrets, information whose leak would put the company in difficulty. Here the hosting question changes shape. It becomes reasonable to run an open model on a machine you control, in your own infrastructure or at a dedicated European host, so that no data leaves your perimeter.

That last option deserves to be known, because people assume it is out of reach for a company of this size. It no longer is. Recent open models, run on a properly sized machine, are good enough for many classification, extraction and summarisation tasks. They remain behind the best proprietary models on complex reasoning, and they cost more in infrastructure than they save in metered calls. It is a trade-off, not an obvious choice, and it is decided case by case.

What the GDPR actually requires

Without going into legal detail, five obligations come up in every project we run.

You must be able to point to a lawful basis for the processing, most often your legitimate interest or the performance of a contract. You must enter the processing in your record, which takes a few lines and is the first thing your national data protection authority will ask for in an inspection. You must sign a processing agreement with every supplier handling personal data on your behalf. You must inform the people concerned, including your own employees when the tool processes their data. And you must apply minimisation, meaning you send the model only what the task needs.

That last point is the most effective in practice and the least applied. Plenty of tasks work just as well on reduced or pseudonymised data. Stripping out names and identifiers before the call to the model, when the task does not require them, settles much of the subject by design rather than by contract.

A word finally on the European regulation on artificial intelligence. The EU AI Act sorts uses by level of risk and places the heaviest obligations on suppliers of high-risk systems, typically in recruitment, credit and personal safety. For ordinary business use, drafting assistance, summarisation, document classification, the main obligations concern transparency: stating that content was generated, informing people who interact with an automated system, and keeping an inventory of your uses up to date. If you touch recruitment or the assessment of individuals, get help, because the regime there is markedly more demanding.

Your checklist

To run through before putting any use involving real data into service.

  • The sensitivity level of the data concerned is identified and written down.
  • The tool in use is a professional offer under contract, not a personal account.
  • The commitment not to reuse data for training is in the contract.
  • The hosting region is known and acceptable given the sensitivity level.
  • The retention period is defined, short, and genuinely applied.
  • The full chain of subprocessors is documented, model supplier included.
  • The processing is in the record and the people concerned have been informed.
  • Only the data the task needs is transmitted.
  • A written rule, five lines long, tells the team what is allowed and what is not.
  • An approved alternative exists for every prohibited use, failing which the rule will be bypassed.

This guide is not legal advice. It describes the points we deal with in every project, and it does not replace the view of your own adviser or data protection officer, particularly on level 3 and level 4 processing.

Our own commitments on handling client data are set out on the data privacy page. And if you want to know which uses are genuinely workable in your company given your sensitivity levels, that is one of the questions a diagnostic settles, area by area.

Next step

Where would you start?

Ten days to map where your teams lose time, put a figure on each lever and name the first system to build. If AI is not the right answer, I will tell you so.

Diagnostic

Étape 1

10 days

€9,500 excl. VAT · fixed fee, one area of the business

  • A map of where the time actually goes in the area audited
  • Each lever costed in hours freed, ranked by effort and by gain
  • The data note: what leaves your systems, and where it is processed
  • A three-month plan, naming the first system to build