Article 50 of the EU AI Act requires an AI system to disclose that a person is talking to a machine. It says nothing about what happens to that conversation once the disclosure is made.
Founders building products on top of a foundation model focus on accuracy and price. The question of what happens to their users’ data once it leaves their product and is processed by a third-party model rarely comes up during vendor selection.
The rule became enforceable on August 2, 2026, and fines for noncompliance run up to €15 million or 3% of global annual turnover, whichever is higher. Retention and training sit completely outside its scope. If your application sends user conversations to a third-party model, that provider’s data practices become part of your product, whether you intended them to or not.
The Default Across the Industry Is Training, Not Asking
Researchers at Stanford’s Institute for Human-Centered AI reviewed the privacy policies of six major US AI developers, including Amazon, Anthropic, Google, Meta, Microsoft, and OpenAI, and found that every one of them used customer chat data by default to train their models, with some retaining that data indefinitely. Lead researcher Jennifer King was blunt when asked whether users of AI chat systems should worry about their privacy, answering “absolutely yes.” Training should therefore be treated as the default assumption unless a provider’s documentation says otherwise. The details, however, vary from one provider to another and deserve close attention before committing to a platform.
Once lauded for its privacy-conscious approach, Anthropic changed its consumer terms so that conversations with Claude are now used for training by default, with users able to opt out. This illustrates how quickly a vendor’s training policy can change, even when privacy was part of its original pitch.
In other models, opting out doesn’t fully work. Google keeps Gemini conversations reviewed by human annotators for up to three years, and turning off the Gemini Apps Activity setting doesn’t undo that. Once a conversation has been flagged for human review, it stays in Google’s systems on a separate retention track regardless of what the user does afterward.
Character.AI trained on user conversations by default and offered an opt-out only to users in the EEA and UK, which puts it at the permissive end of the data practices covered here. Separately, the company has faced legal and reputational challenges. In January 2026, Character.AI and Google agreed in principle to settle a group of lawsuits alleging harm from the platform’s chatbot interactions with minors, unrelated to its data-training policy. Google was named as a co-defendant because of its ties to Character.AI’s founders, and while no liability was admitted and the terms weren’t disclosed, the case shows that legal exposure tied to an AI vendor can reach the companies that back it.
Retention windows, human review triggers, and subcontractor access rarely show up outside a privacy policy’s fine print. Companies integrating these APIs frequently have to investigate thoroughly to find out how long user data sits in review or who else can see it. Almost none of that detail ever reaches end users, which means a company can inherit a vendor’s data practices and pass the same exposure on without either side realizing it happened.
Why This Is a Founder Problem and Not Just a Policy Debate
If your product routes user conversations through a foundation model with default training enabled, you’ve accepted a data policy on your users’ behalf without telling them, even if you never intended to. Article 50 doesn’t require you to disclose that decision because it only covers the fact that they’re talking to AI.
Third-party access widens that exposure further. Model providers regularly outsource data annotation and safety review to subcontractors. In May 2026, Texas Attorney General Ken Paxton opened an investigation into Meta’s AI glasses after data annotators at a Kenya-based subcontractor called Sama reportedly said they could access footage of users’ private moments, including bathroom visits. Meta disputed that this reflected the full picture of its practices, and the investigation is ongoing. While the example involves wearable cameras, the same issue can apply to conversational AI. A promise that a company doesn’t sell your data says nothing about who else can read it internally before it ever reaches a training pipeline.
What to Check Before You Pick a Vendor
Before connecting a foundation model to real user data, there are several questions every founder should answer.
The data processing agreement is a place to start. That agreement is the document that governs how your users’ data is handled, and it’s worth reading in full.
The contract itself deserves as much attention. Vendors often offer stricter no-training terms on enterprise or API contracts than on consumer chat products, so confirm in writing which terms apply to your specific integration. Many companies often make the wrong assumption that the enterprise-grade privacy language covers them by default.
Opt-out provisions are also worth examining. Google’s Gemini settings illustrate why, since in that model, disabling activity tracking doesn’t stop human review from happening. A vendor should be able to say plainly whether an opt-out halts training, human review, or both.
Finally, map where the data goes downstream of the vendor. Annotation and safety review are frequently outsourced, and a vendor’s own privacy promises don’t automatically extend to every subcontractor in that chain.
None of these questions should be treated as a one-time exercise. Vendor terms change, sometimes quickly, and the vendor you signed with last year may not be operating under the same rules today.
The Market Can Set a Standard the Regulation Hasn’t Reached Yet
Article 50 is unlikely to be the last word on AI data governance from the EU, and founders shouldn’t wait for the next round of regulation to catch up with the gap it left behind. Companies can get ahead of that gap now by building an explicit and visible opt-in for training on their own users’ data. There’s a real line between improving a model and quietly cashing in on user trust. Plenty of AI services still cross it by using conversations to train without ever asking for explicit, informed consent. Founders who draw that line for themselves now won’t need to retrofit compliance later.
It’s a move with precedent. Encrypted messaging apps and password managers didn’t wait for regulators to make strong defaults standard. They made those choices on their own and turned them into a market position. AI products have the same opportunity. The founders who draw that line before it’s mandatory are the ones who will be trusted with the next sensitive conversation their users have.

