Skip to main content

The Paste Test: What Your Team Can Safely Put Into AI

Is it safe to paste company data into AI?

The commonest way an organisation loses control of its information is not an attack. It is someone pasting a document into a chat box to get help with it, at four in the afternoon, to hit a deadline. The useful question is not whether people will do that. It is which categories genuinely must not be pasted, and which are fine.

6 of the 9 categories below should never go into a general assistant. Some are fine on the right tier. One should be actively encouraged, because a policy that forbids everything is the reason people use their personal accounts instead.

What are you about to paste?

Tick everything the text contains, including the parts you were not thinking about. Nothing is stored or sent anywhere.

Pick at least one to see the answer.

All 9 categories

Credentials

Do not paste

Passwords, API keys, tokens, connection strings

A credential pasted into any chat box must be treated as disclosed, immediately and permanently. It is now in a log, a prompt history, a support tool and possibly a training set, and no retention promise can un-know it. This is the one category where there is nothing to weigh up.

Not primarily a privacy question. It is an access control failure, and in a regulated environment it is a reportable one. If the credential reached anything holding personal data, the clock on a breach assessment has already started.

The rule to write down. Credentials never go into an AI tool, in any circumstance, including to ask what is wrong with them. Rotate anything that already has.

Special category data

Do not paste

Health, biometric, genetic, sexual orientation, religion, union membership, criminal matters

This is the data class where a mistake is not recoverable by apology. Health records and biometric identifiers pasted for summarising or drafting are being processed by a new processor, in a new place, usually with no record that it happened.

Under the GDPR, Article 9 prohibits processing special category data unless a specific condition applies, and legitimate interests is not one of them. There must also be a lawful basis under Article 6, a record of the processing, and a controller to processor contract with the AI provider under Article 28. An employee pasting into a consumer tool has none of those.

The rule to write down. Special category data goes only into a named system approved for it in writing, never into a general assistant, and never as an example.

Customer personal data

Do not paste

Customer names, emails, addresses, order or account records

Pasting a customer list to reformat it is processing personal data through a third party. It is lawful only if the organisation has decided it is, in advance, and can show it. Most cannot, because the decision was never made, it was improvised by whoever had the deadline.

The GDPR needs a lawful basis, transparency to the data subject, a processor contract, and a transfer mechanism if the provider processes outside the EEA or UK. The failure is rarely the paste itself. It is that no record exists, which turns an arguable decision into an indefensible one.

The rule to write down. Personal data goes only into tools already covered by a signed processor agreement, and the list of those tools is short, written down, and known.

Material under NDA

Do not paste

Contracts, term sheets, or anything received under an NDA

Disclosure to a third party is disclosure, whether the third party is a person or a service. Most non-disclosure agreements permit sharing only with named categories of recipient, and a general AI provider is not usually one of them.

Contractual rather than regulatory, which makes it easier to overlook and no less expensive. The counterparty does not need to prove damage to enforce a breach, and the disclosure is usually discoverable from your own logs.

The rule to write down. Nothing received under an NDA is pasted anywhere until someone has read what the NDA permits. If nobody has read it, the answer is no.

Material non-public information

Do not paste

Unreleased results, forecasts, deal terms, anything price sensitive

Draft numbers pasted in for a cleaner summary have left the perimeter before they were announced. For a listed company or anyone advising one, the exposure is not embarrassment, it is a regulator asking who had access and when.

Market abuse and insider dealing regimes require issuers to control inside information and keep lists of who holds it. A tool that received the numbers and cannot be listed is a hole in that control.

The rule to write down. Price sensitive material stays inside approved systems until it is public. After announcement, the same text is usually fine.

HR and candidate data

Do not paste

Employee records, candidate CVs, performance notes, grievances

The most frequently pasted sensitive category, because summarising CVs and drafting reviews is exactly what these tools are good at. It is also the category where the person affected has the strongest interest and the least visibility.

Personal data, so everything under customer data applies. In addition, using AI to assess candidates or employees engages employment rules directly: Illinois requires notice, explanation and consent for AI analysis of video interviews and prohibits AI with a discriminatory effect in employment, Maryland requires consent for facial recognition in interviews, and New York City requires an annual independent bias audit of automated employment decision tools. Our own answer page covers what is and is not required, including the fact that no US law requires you to train staff on any of it.

The rule to write down. HR data goes only into approved systems, and any AI that influences a hiring or performance decision is a decision the organisation must be able to explain and defend.

Source code

Approved tools only

Source code, infrastructure config, internal architecture

Usually acceptable and occasionally catastrophic, and the difference is what the code contains. Code carrying embedded credentials, customer data in fixtures, or the specific mechanism that is your actual advantage is not the same as a utility function.

Trade secret protection depends on reasonable steps to keep the thing secret. Disclosure to a provider whose terms permit training can undermine that status, which is a loss you cannot buy back. Check whether your enterprise tier trains on your input, because the consumer tier often does and the enterprise tier often does not.

The rule to write down. Code is fine in an approved tool on a no-training tier. Scan for secrets before pasting, and keep the crown jewels out regardless of tier.

Internal strategy

Approved tools only

Internal strategy, board material, restructuring plans

Rarely a legal problem and often a commercial one. The realistic risk is not the provider reading it, it is the same document arriving in a support ticket, a prompt log, or a screenshot in a chat with someone who should not have it.

Mostly governance rather than statute, unless it touches people or markets, in which case the categories above take over. The question a board will ask afterwards is whether a policy existed, not whether the tool was good.

The rule to write down. Approved tools only, and remember that anything about a named employee has become personal data whatever the document is called.

Already public

Go ahead

Published material, public filings, your own marketing copy

Nothing that has not already happened. Public material is public, and refusing to let people use AI on it is the fastest way to make them use it on everything else somewhere you cannot see.

None from disclosure. Ordinary duties still attach to what you do with the output, particularly accuracy and, where the output is shown to people, the transparency expectations that now apply to synthetic content.

The rule to write down. Encourage this. A policy that only forbids gives people no permitted path, and an unusable policy is the reason shadow AI exists.

The five questions that settle the rest

We name no vendors and rate none, because retention and training behaviour changes by product, by tier, by contract and by month, so any table we published would be wrong before it was useful. These five answers determine what your policy can safely permit, and they live in an agreement rather than on a marketing page.

  1. 1. Does this tier train on our input, and is that in the contract or only on a web page?

    A web page can change on a Tuesday. The commitment that matters is the one in the agreement, and the consumer and enterprise tiers of the same product frequently differ on exactly this.

  2. 2. How long is our input retained, and where?

    Retention drives everything downstream: deletion rights, breach scope, and whether an auditor can be told the data is gone. Zero retention is a real offering and it is usually not the default.

  3. 3. Which subprocessors see it, and in which countries?

    This is the question that decides whether an international transfer is happening. It is also the one most likely to be answered vaguely, which is itself an answer.

  4. 4. Can a human at the provider read our prompts, and under what circumstances?

    Abuse review and support access are legitimate and normal. They are also processing, and a policy written as though only machines see the data is wrong.

  5. 5. What do we get if there is an incident?

    Notification timelines and audit rights are ordinary in a processor agreement. If they are absent, the tool is not ready for anything on the do-not-paste list.

Why banning everything makes it worse

Shadow AI is not defiance. It is what happens when the sanctioned path is missing, slow, or forbids the work people actually have to do, so they use a personal account where you cannot see anything at all. A policy that names a permitted tool, a permitted tier and a short list of categories that never go in beats a policy that says no, because the second one is not followed and produces no record that anything was decided.

Questions people ask

Is it safe to paste company data into ChatGPT or another AI tool?

It depends entirely on the category of data and the tier you are on, not on the tool's reputation. Credentials, special category data, customer personal data, material under NDA, price sensitive information and HR records should not be pasted into a general assistant at all. Source code and internal strategy are defensible on an approved enterprise tier with a no-training commitment in the contract. Already public material is fine and should be encouraged, because a policy with no permitted path is the reason people go around it.

What is shadow AI?

Employees using AI tools the organisation has not approved, usually on personal accounts, to do work faster. It is not defiance. It happens when the sanctioned path is missing, slow, or forbids everything, so the practical fix is providing a good approved tool rather than issuing a stricter ban, which mainly moves the activity somewhere you cannot see.

Does pasting personal data into an AI tool breach the GDPR?

It can, and the failure is usually not the paste itself. Processing personal data through a third party needs a lawful basis, transparency to the data subject, a controller to processor contract under Article 28, and a transfer mechanism if processing happens outside the EEA or UK. Special category data under Article 9, such as health or biometric data, needs a specific condition on top of that. An improvised paste into a consumer tool has none of these and, critically, leaves no record that a decision was ever made.

Can AI vendors train on our confidential data?

Some tiers do and some do not, and the honest answer for any specific product is in your contract rather than on a marketing page. Ask whether the commitment not to train is contractual, how long input is retained, which subprocessors see it and in which countries, and whether a human at the provider can read prompts. Those four answers determine what your policy can safely permit.

Does pasting source code into AI lose trade secret protection?

It can weaken it. Trade secret protection depends on taking reasonable steps to keep the information secret, and disclosure to a provider whose terms permit training is hard to characterise as a reasonable step. On an enterprise tier with a contractual no-training commitment the position is much stronger. The practical rule is to scan for embedded credentials first, use an approved tier, and keep the specific mechanism that is your actual advantage out of it regardless.

Related, and free

Whether any law requires you to train staff on this answers the question your board will ask next, including the part where the honest answer is no. The cross-compliance graph shows where the GDPR, DORA, NIS2 and the EU AI Act land on the same obligation.

This page is information, not legal advice, and nothing you tick is stored or sent. Where the answer depends on your contract with a provider we say so rather than guessing, which is why the five questions exist instead of a vendor table.