In April, someone asked a chatbot for help drafting a blog post about cloud security. Unpublished. With details of a corporate project in it. When they were done they pressed the share button, probably to show a colleague.
That document ended up reachable from Google.
There was no attack. No vulnerability was exploited. Nobody stole any credentials. The person who exposed the document was the person who wrote it, and at no point did they feel like they were publishing anything. This is what shadow AI actually looks like, and it is considerably less cinematic than an intruder on the network.
What happened, and what did not
Over the weekend of 25 July, Reddit users found that conversations people had shared from Claude, Anthropic's chatbot, were showing up in search results. Shortly afterwards the same proved true of Artifacts, the apps, reports and spreadsheets generated inside the tool. Anthropic fixed it on Monday the 27th and Google began removing the links, though for a while they were still reachable through Bing and Brave. The BBC, Fortune, Axios and much of the Spanish press covered it.
It is worth being precise about what did not happen, because most of the headlines overshot.
This was not a breach. There was no hack. Claude chats are private by default, and the affected links gave nobody access to an account or to anyone's other conversations. Every one of those pages existed because a person chose to create it. Anthropic put it this way in its statement to the BBC:
"When someone shares a conversation, they are making that content publicly accessible, and like other public web content, it may be archived by third-party services."
That is a reasonable answer. And it still does not close the problem, for a reason that has little to do with privacy and a lot to do with expectations: sending a conversation to a colleague is not the same as leaving it indexed in a search engine. The share dialog warns that "anyone with the link" can view it. It does not say the link may end up in Google.
What was inside
The inventory is the uncomfortable part, and it comes from tier-one outlets:
- An unpublished draft on cloud security containing details of a corporate project (BBC)
- CVs with names, contact details and work history (BBC)
- What appeared to be proprietary work research, including healthcare, with transcripts of private conversations (BBC)
- Cryptocurrency wallet keys, names and addresses (Fortune)
- Documents belonging to companies, universities and medical projects (Gizmodo)
On volume, honesty is required: there is no reliable figure. The BBC counted more than 200 conversations across at least 25 pages of results. Gizmodo saw around 612 pages still in Bing. Other outlets said thousands. Each measured at a different point during the takedown, so any round number you read is a partial snapshot at best.
What matters is not how many there were. It is what kind.
The technical detail almost nobody reported
Most of the coverage settles this with the phrase "misconfiguration" and moves on. It is worth stopping, because this is the part that is actually useful to you.
You can check it yourself. The file at claude.ai/robots.txt, as of today, contains this line:
Disallow: /share/*
So shared conversations were blocked to crawlers. And according to Forbes they already were in September 2025, when this happened the first time. They were indexed anyway.
That is not a contradiction. It is exactly how this works. robots.txt is a crawl directive, not an indexing directive. Google's own documentation says so without ambiguity:
"A page that's disallowed in robots.txt can still be indexed if linked to from other sites."
In plain terms: if someone pastes that link into a forum, a social network or any public page, the search engine learns the URL exists and may list it, even though it promised not to crawl the contents. To actually keep an address out of results you need something else: a noindex tag, authentication, or removing the page. The affected pages carried no noindex.
There is a second verifiable detail. That same robots.txt contains no rule at all for the /public/ path, which is where Artifacts live. The generated apps and documents were not covered even on paper.
Now the question you should be asking: how many of your organization's controls are of this kind? How many times has someone said "that's blocked in the robots file" and everyone nodded? Checking that on your own domain takes two minutes and is probably the highest-return thing you do this week.
This is not the first time. It is the fourth.
| Date | Product | Reported scope |
|---|---|---|
| August 2025 | ChatGPT | Thousands of conversations indexed (Fast Company) |
| August 2025 | Grok | More than 370,000 conversations (Forbes) |
| September 2025 | Claude | Google estimated just under 600 (Forbes) |
| July 2026 | Claude | Conversations and Artifacts |
Four episodes in a year, three different providers, and the second for the same one. In September 2025 Anthropic explained that it blocked Google's crawlers. That was true, and it still is. Ten months later it happens again, this time extended to a path that rule never covered.
The counter-example is worth naming. When ChatGPT had its episode in August 2025, OpenAI did not adjust a configuration file: it removed the feature that let shared chats be discoverable, described it as a "short-lived experiment", and worked with search engines to de-index. It changed the product, not the robots file.
As Fortune puts it, the fact that OpenAI, Anthropic and xAI have each stumbled into their own version of the same problem suggests it is something the industry has struggled to fix permanently.
That is the operational conclusion. This is not about picking the provider who will not get it wrong.
Why this is not Anthropic's problem but yours
This is where most of the analysis reaches for the wrong regulation.
A great deal has been written this week about the EU AI Act, and it is true that on 2 August 2026, days from now, the Commission and the AI Office gain supervision powers over general-purpose AI model providers, with fines of up to 3% of worldwide annual turnover or 15 million euros. But those obligations point at Anthropic, OpenAI and xAI. Not at the company whose employee used the chatbot.
What exposes you is the GDPR.
If someone on your team pastes customer, patient or candidate data into a chatbot and shares a link that becomes publicly accessible, your company is the controller and that is an unauthorized disclosure: a personal data breach in the sense of Article 4(12). From there Article 33 opens, giving you 72 hours to notify your supervisory authority once you become aware, and Article 34 requires telling the affected individuals where the risk is high. Fines reach 4% of global turnover.
And if what was shared included healthcare transcripts, you are in Article 9, special category data, with a considerably higher risk threshold.
None of this depends on the provider getting it right. It depends on what left your organization.
The context does not help either. According to IBM's Cost of a Data Breach 2025 report, one in five organizations that suffered a breach had shadow AI involved, and those breaches cost roughly 670,000 dollars more than the rest. 63% of affected organizations had no AI governance policies. Among those reporting AI-related incidents, 97% acknowledged they lacked adequate access controls.
What to do this week
Five concrete things, in order of effort.
- Review your own shared links. In Claude it is under Settings, Privacy, Shared chats: you can switch a conversation back to private and disable its link. Note that revoking stops it opening but does not delete copies already archived or forwarded. Anthropic says as much itself.
- Inventory which AI tools your people actually use, not which ones are approved. They are different lists, and the gap between them is the problem.
- Write into your AI policy that sharing is publishing. In those words. Most people do not read a private link as a web page, and that is not their fault.
- Get personal and customer data out of consumer-tier tools. Real work goes through enterprise agreements, with no-training and retention terms. For everything else, strip the data before pasting: there is a free tool below that does exactly that.
- Audit your own domains.
robots.txtis notnoindex. Check what of yours is indexed that should not be.
And one thing to settle in advance: decide who starts the 72-hour clock. When this happens, time runs from when you become aware, not from when you finish investigating.
A free tool for the part that does have a quick fix
There is a part of this problem that takes thirty seconds to solve, and it is the part that comes up most: the personal data people paste into a chatbot without thinking. Customer names, an ID number, a phone number, an IBAN, an email address inside an email they copied wholesale.
That is why we have published the Moviwa anonymizer, free and with no signup.
It works like this. You paste the text, or upload a .txt, .md, .pdf or .docx. The tool detects the personal data and replaces it with tokens like [PERSON_1] or [ID_1]. You copy that cleaned text and use it in ChatGPT, Claude, Gemini or whichever you prefer. When the AI answers, you paste the response back and recover the original values in one click.
The important part: everything happens inside your browser. The document is not uploaded to any server, including ours. It runs two detection engines, one pattern-based (national ID numbers, phone numbers, emails, IBANs, cards, licence plates) and a local AI model for names that downloads once and runs entirely on your machine. The token map lives only in your browser's local storage, expires after seven days, and you can wipe it whenever you want. No accounts, and no analytics on the contents of your documents.
Now the honest part, which is directly relevant here: this does not stop anyone sharing a conversation. What it does is ensure that if that conversation reaches a search engine, what is readable is tokens rather than your customer's ID number. It also does not protect your company's own knowledge, only personal data, so the cloud architecture draft from the opening would still be legible, because the problem there was not personal data but the content itself. That takes policy and judgement, not a tool.
It covers the frequent, avoidable part. It does not cover the part that requires deciding what may be pasted at all.
What is left
A conversation with a chatbot gives a very convincing sense of intimacy. It is a window, one to one, with nobody watching. That feeling holds right up to the moment someone presses share, and from then on the privacy of what you wrote depends on a search engine's configuration and on nobody pasting the link anywhere.
The governance question is not whether your provider will get it right next time.
It is what you have already handed to a share button.
Start with the easy part: try the anonymizer. It is free, needs no signup, and runs entirely in your browser. Pass it to whoever on your team handles customer documents every day.
Then the rest. The anonymizer solves one specific case for one person. Moviwa protects the whole organization: AI-based name detection, documents with their original formatting preserved (PDF, Word, Excel), a browser extension that is transparent to the user, full auditing and GDPR compliance. If you want to see what is being pasted into a chatbot from your company today, request a demo and we will look at your case.
