Answers

Does ChatGPT train on data employees paste into it?

It depends on the account tier. Consumer tiers have historically used conversations for model improvement by default, with an opt-out; business and enterprise tiers contractually exclude training. Verify the current terms directly — and note that training is only one of the risks.

ONE PASTE SANCTIONED PERSONAL
Same hostname. Same interface. Entirely different legal position.
Also in answers Prevention Legality Control overlap Architecture

The tier distinction

That is the whole answer.

Your policy applies to the tenant you bought. Your employees’ behaviour applies to whatever they are logged into — and someone with your enterprise seat and a personal free account in the same browser can move between them without noticing.

Nothing in the enterprise agreement reaches the personal session.

Training posture by account type
Account typeTypical training postureWho is using it
Free / personal paidUsed for model improvement by default; opt-out availableAnyone who signed up with their work email in five minutes
Team / BusinessExcluded by contractWhoever procurement onboarded
Enterprise / API with a DPAExcluded by contract, retention negotiableYour sanctioned deployment

Verify before relying on this. Provider terms change, sometimes with little notice. Any page that states another company’s data-handling terms as settled fact — including this one — should be checked against the live agreement during procurement.

The larger risks

Five things that survive a no-training clause.

Training dominates the conversation because it is vivid. These are less discussed and, for most organisations, larger.

  1. 01

    Retention

    Conversations persist for a defined window even when excluded from training. Data in a third party’s storage is data in scope for their breach.

  2. 02

    Human review

    Most providers reserve the right to have staff review flagged conversations for abuse and safety monitoring. That is a legitimate safety control, and it is also a human reading your content.

  3. 03

    Subprocessors

    The provider’s infrastructure and support vendors sit inside the trust boundary you accepted.

  4. 04

    Cross-border transfer

    Where processing happens has GDPR consequences independent of whether a model was trained.

  5. 05

    Your own obligations

    If the data was PHI, submitting it to a vendor without a BAA is a disclosure regardless of what the vendor does next.

Deletion

“We’ll just delete the conversation.”

Deleting a conversation removes it from the interface and, after the retention window, from storage. If the content was already used in training, deletion does not extract its influence from the model’s weights — that is not a capability anyone currently offers at production scale.

Which is the real point: the effective control is at submission time. Once sensitive data has been sent to a third party, every remaining option is mitigation of a disclosure that has already occurred.

Straight answers

Do business and enterprise AI tiers train on submitted data?

Major providers contractually exclude business and enterprise API and workspace traffic from model training. Terms change, so treat any statement of this — including this page — as something to verify against the provider’s current agreement during procurement.

If training is excluded, is the data safe?

Not entirely. Retention windows, human review for abuse monitoring, subprocessor access, breach exposure and cross-border transfer all persist independently of training. Training is the most discussed risk, not the largest.

Can data already submitted be removed from a trained model?

Practically, no. Deleting a conversation removes it from the interface and eventually from storage under the provider’s retention policy. Influence already absorbed into model weights is not something deletion addresses.

Does an opt-out apply retroactively?

No. Opting out changes how future submissions are handled. Content submitted before the opt-out was governed by the terms in force at the time.

Stop it at submission instead.

If the data never leaves the device, what the provider does with it stops being your question.