Is it safe to paste client data into ChatGPT?
What ChatGPT, Claude, and Copilot promise about your data, what the Privacy Act requires, and the honest answer for professionals handling client files.
A bookkeeper pastes a client’s tax file number and three years of transactions into ChatGPT and asks for a tidy summary. A lawyer pastes a client’s name and matter details into Claude to check a citation. A psychologist pastes session notes into Copilot to draft a letter. None of them think they are doing anything unusual. All three have just run an experiment they have not read the terms of.
The honest answer to “is it safe” is not yes or no. It depends on which product you are actually using, and the Privacy Act has an opinion regardless of which one that is.
What ChatGPT actually promises
OpenAI’s own position is clear and worth reading in full rather than trusting a paraphrase: on the free, Plus, and Pro consumer plans, ChatGPT uses your conversations to train its models by default, unless you turn that off yourself in Settings under Data Controls (vendor policy, OpenAI Help Center). Once a conversation has been included in a training run, it cannot be pulled back out.
ChatGPT Business, Enterprise, Edu, for Healthcare, for Teachers, and the API platform sit on different terms: by default, none of them use your inputs or outputs for training (vendor policy, OpenAI, “Business data privacy, security, and compliance”). That is the tier built for organisations, provisioned by an admin, usually with a Data Processing Addendum behind it.
Most professionals typing into chatgpt.com from a personal login are not on that tier. They are on the one that trains by default.
What Claude and Copilot promise
Anthropic draws the same line in a different place. Consumer Claude, Free, Pro, and Max, retains your chats for 30 days by default. If you opt into “Model Improvement” in your privacy settings, your chats and coding sessions can be retained in de-identified form for up to five years and used in training (vendor policy, Anthropic Privacy Center). Claude for Work, Claude Enterprise, and the API sit under separate commercial terms: not used for training, full stop, no toggle required.
Microsoft’s commercial tier makes a similar commitment. Prompts, responses, and the Microsoft Graph data behind them are not used to train the underlying foundation models for Microsoft 365 Copilot, provided it is deployed under a commercial licence and covered by the same contractual terms that already apply to a business’s email and SharePoint files (vendor policy, Microsoft Learn, “Data, Privacy, and Security for Microsoft 365 Copilot”). The free, personal Copilot a person signs into from home is a different product on different terms.
The pattern holds across all three vendors. The paid, admin-provisioned, business tier gets a no-training commitment in writing. The free tier a professional opens from a personal account, on a Tuesday, to get through the afternoon, usually does not.
What the Privacy Act actually requires
None of the above answers the question an Australian professional actually needs answered, because a vendor’s training policy is not the same thing as compliance with the Privacy Act 1988. The Act and the Australian Privacy Principles (APPs) apply to any personal information handled through an AI system, including where that information is only used, not trained on (regulator, OAIC, “Guidance on privacy and the use of commercially available AI products”, published October 2024, updated January 2025).
The relevant principle for pasting client data into a chatbot is APP 6: personal information can only be used or disclosed for the purpose it was collected for, unless the client consented or would reasonably have expected the secondary use. The OAIC’s own worked example is close to home: an insurance company’s staff paste a customer’s claim details, including sensitive health information, into a public chatbot to draft an assessment report. The OAIC’s reading is that this is very likely a disclosure of personal information to the chatbot’s owner, for a purpose the customer was never told about at the time their information was collected, and most businesses will not be able to show it falls inside an APP 6 exception.
The OAIC states its own best-practice position plainly: “the OAIC recommends that organisations do not enter personal information, and particularly sensitive information, into publicly available generative AI tools, due to the significant and complex privacy risks involved” (regulator, OAIC guidance, above). It also notes that once personal information is in a generative AI system, it is “very difficult to track or control how it is used, and potentially impossible to remove.”
Two more principles matter in the same guidance, briefly. APP 3 treats AI-generated inferences about a real person, including hallucinated ones, as a fresh collection of personal information in their own right. APP 8 governs disclosure overseas, relevant because none of the three consumer products above processes Australian client data on Australian soil by default. None of this changes because a vendor’s marketing page says “we take privacy seriously.” The OAIC’s guidance is explicit that a policy commitment from the provider does not discharge a business’s own obligations under the Act.
Where this actually bites
This is not an abstract compliance exercise for most of the professions handling sensitive client material day to day:
- Lawyers working under legal professional privilege, where a client’s name attached to a matter is itself sensitive
- Accountants and bookkeepers handling tax file numbers, bank details, and AML/CTF-covered client identification
- Psychologists and counsellors whose session notes are health information, a category the Privacy Act treats with additional care
- Real estate agents, valuers, and building inspectors handling a client’s financial position alongside personal circumstances
For the AHPRA-registered practitioners in that list, the same client relationship carries a second, unrelated obligation that is easy to miss: their public website is advertising under the National Law, with its own set of prohibitions. What AHPRA’s advertising rules mean for your website covers that regime.
In each case, the value of AI is real. It drafts faster, checks patterns a tired reader misses, and analyses volumes of text a person would not get through in a working day. The barrier is not the AI. It is that the professional’s client never consented to their name and file appearing in a prompt sent to a third-party server, and telling them after the fact is not the same as asking first.
The honest limit
None of the paid, no-training tiers described above are an architectural guarantee, and it would be dishonest to present them as one. Even under the strongest commercial terms, the data still leaves the professional’s device in identifiable form and exists, at least briefly, on infrastructure they do not own and cannot inspect. A contract is a promise about how that data will be handled. It is not a mechanism that prevents the data from being there in the first place. What happens to your data when you press ‘Send’ on an AI tool sets out that fuller spectrum, from no protection through to hardware-secured processing, and where policy commitments sit on it.
Masking a client’s name before it ever reaches the AI is a stronger position than trusting a training policy, but it is not a compliance guarantee either, and no automated tool should be sold as one. Even well-masked text can sometimes be re-identified by a knowledgeable reader from context alone, industry, revenue, family detail, if enough of it survives the strip. Nothing in this essay, or in any tool built on this approach, constitutes legal advice about a specific practice’s Privacy Act obligations.
The practical answer
The most reliable way to close the gap between “the vendor promises not to train on this” and “my client never agreed to this leaving my device” is to make sure it never leaves identifiable. De-identify is a free tool that does exactly that: it strips names, organisations, and places using an in-browser named-entity model, and Australian identifiers, phone numbers, emails, addresses, Medicare numbers, ABNs, TFNs, BSBs, and card numbers, using pattern matching, all inside the browser before anything is pasted anywhere. Nothing is uploaded. That is checkable, not just claimed: open your browser’s developer tools, watch the network tab, and mask something yourself. The full architecture behind it, including how a coaching platform built the same approach into production, is described in How we built on-device de-identification so AI never sees real names.
Masking first changes the APP 6 analysis in a professional’s favour: if a client’s name and identifiers never leave the device, there is a real argument that no personal information has been disclosed to the AI provider for that input at all, only tokens. That argument still depends on managing re-identification risk sensibly, and it is worth treating as one part of a considered approach rather than the whole of one.
Getting that considered approach right, across a whole practice rather than one paste box at a time, is what an AI strategy and governance engagement is for. That is the conversation we have with clients who are further along than “can I paste this in” and need a policy their whole team can actually follow: see AI strategy, governance, and training, or start with a conversation.