Repeatedly, publicly, and at organisations with better security programmes than most.
The documented cases fall into three groups that get muddled together constantly, and they carry completely different lessons.
The ordinary case, and the one this site exists for. No attacker, no vulnerability, no failure of any control. A competent person trying to finish a task faster.
Within weeks of the semiconductor division being permitted to use ChatGPT, separate incidents were reported in which employees entered internal source code and the contents of a recorded internal meeting into the tool. The company subsequently restricted generative AI use on company devices while it built its own internal alternative.
Reported by The Economist Korea, spring 2023, and widely covered including by Bloomberg on the subsequent restriction.
An internal warning told employees not to share confidential information — including code — with ChatGPT, noting that some responses appeared to resemble internal material.
Reported by Business Insider, January 2023.
The most-cited figures come from Cyberhaven, whose product monitors exactly this, and who reported that a meaningful share of workers had pasted company data into ChatGPT, some of it confidential.
That is vendor research, and it should be read as such. The company sells a product that detects the problem it is measuring, and its sample is organisations that had already bought monitoring. The direction is almost certainly right; the specific percentages are not the kind of number to put in a board paper without saying where it came from.
Different lesson entirely. Here nobody did anything wrong at all, and the data moved anyway.
A bug in an open-source library used for caching allowed some users to see titles from other users' conversation history, and exposed partial payment information belonging to a small percentage of subscribers during a several-hour window. OpenAI took the service offline, published a postmortem describing the cause, and notified affected users.
Disclosed by OpenAI in its own incident postmortem, March 2023.
Worth being fair about: the disclosure was prompt, specific and technically candid, which is better behaviour than most organisations manage. The point is not that the vendor was careless. The point is that "we don't train on your data" and "your data cannot be exposed" are different promises, and only one of them was ever made.
Security researchers reported finding a database belonging to the AI provider DeepSeek accessible over the internet without authentication, containing chat history along with internal keys and operational metadata. It was secured after disclosure.
Reported by Wiz Research, January 2025.
An ordinary cloud misconfiguration, of the kind that happens to every category of company. It matters here only because of what was inside: the conversations. Every AI provider is also a database operator, and the fastest-moving ones are operating at a scale their security programme is still catching up with.
Italy's data protection authority ordered a temporary halt to ChatGPT's processing of Italian users' data in March 2023 over lawful basis and transparency concerns, and the service returned after changes. The same authority issued a substantial fine in December 2024.
Provvedimenti of the Garante per la protezione dei dati personali, March 2023 and December 2024.
The relevance to you is not the fine. It is that a regulator examined how prompts are processed and found the question worth ruling on — which means the assumption that a prompt is a private message between you and a machine has already been formally tested and found wanting.
A page like this is only useful if it declines to overstate, so here is the strongest argument against its own thesis.
There is no well-documented public case of a model reciting one organisation's pasted secret verbatim to an unrelated user. That specific fear — you paste your API key, a competitor asks the right question and gets it back — is the one people hold and the one with the least evidence behind it. Anyone telling you it happens routinely is going beyond what is known.
The real risk is duller and much more likely: your text left your organisation, it is stored on infrastructure you don't control, some humans may read it for abuse monitoring, it is subject to subpoena and to the vendor's own breach exposure, and you disclosed someone else's data to a processor you never authorised. None of that requires the model to leak anything at all.
A paste leaves no trace on your side. No alert, no log entry, no ticket, nothing to review. If it went through a browser on a personal account, your own monitoring never saw it and never will.
Every incident on this page became public because of something incidental — an internal memo that leaked, a vendor's own disclosure, a researcher scanning the internet. None of them surfaced because the organisation detected the disclosure and reported it.
Which means the honest reading of "we haven't had one of these" is "we have no mechanism that would tell us." The absence of detection is not the absence of disclosure, and the gap between those two is where this entire problem lives.
Not "ban AI" — that policy has been tried, and it produces people using their phones instead, which is the same disclosure with none of the visibility.
What works is narrower: know which five categories genuinely must not go in, make checking easy enough that people actually do it, and make reporting a mistake safe enough that you hear about it on the day rather than in a regulator's letter.