Back to Insights
Sovereign Intelligence

Your Team Is Pasting Company Secrets Into AI Right Now

Terry LyonSeptember 15, 2026

Someone on your team pasted a customer list, a pricing sheet, or a draft contract into a chatbot this week. They were well-intended in trying to get their work done faster. The question that matters now is: where did that data just go, and is it still yours.

Does AI train on your company data?

It depends entirely on which version of the tool your employee used. The free chatbot open in a browser tab and the business or API version of the very same model operate under different rules, and the gap between them is where the exposure lives. On the major consumer chat products, conversations can be used to improve and train future models unless a user goes into settings and turns that off. On the business tiers from the same vendors, the default is reversed. OpenAI, for example, states that data submitted through its API, ChatGPT Enterprise, ChatGPT Team, and ChatGPT Business is not used to train its models, and that training happens only if a customer explicitly opts in.1

Which door did your person walk through? Many companies have not answered it, because the choice got made by default rather than on purpose. Employees signed up for the free tool on their own, with a personal email, and the training default came along for the ride.

Does "we don't train on your data" mean your data is private?

No, because training and inference are two different things a vendor does with your data, and the common promise covers only the first. Training is when your input is used to adjust the model itself, so that what you typed can shape an answer given to someone else next month. Inference is the ordinary act of sending your prompt to the model and getting a response back. When a vendor says it does not train on your business data, it is speaking about the first activity. The inference is not covered by that promise.

Here is what inference-time handling generally looks like on the business tiers today. The major vendors do not train on business or API data by default, and separately they retain your inputs and outputs for a limited period, commonly around 30 days, so their systems can watch for abuse and misuse before the data is deleted. Content that gets flagged can be held longer, and authorized staff or contractors may review it. A true zero-retention setup, where nothing is stored at all, is generally available as an option, but you have to qualify for it and request it rather than receive it by default.2

"We do not train on your data" is a real commitment. It also has a narrower coverage than it sounds.

How often is this actually happening?

Far more often than many leaders would guess, and on a schedule. In its 2026 AI Adoption and Risk Report, the data-security firm Cyberhaven found that 39.7 percent of all interactions employees have with AI tools involve sensitive data, and that the average employee feeds sensitive information into an AI tool roughly once every three days.3 The most exposed categories include source code, internal project data, and regulated information.

Where is this traffic going? A large share of it flows through personal accounts that never touch a company control. Cyberhaven measured the share of usage running through personal rather than corporate accounts at 58.2 percent for one major assistant and 60.9 percent for another.3 This is a regular occurrence and it is happening in a channel your IT team cannot see.

What happens to the data after someone hits delete?

You may not fully control that, even on a business tier, and a recent court case makes the point better than any policy memo could. In the copyright suit The New York Times brought against OpenAI, a federal magistrate judge in May 2025 ordered the company to preserve and segregate output logs that would otherwise have been deleted, including chats users had explicitly deleted and data that privacy law would ordinarily require the company to erase. A district judge affirmed the order in June 2025, and although its scope was later narrowed, data preserved under it stayed preserved.4

Conversations employees believed were deleted were held anyway, because a litigation matter that you were not party to reached in and froze the delete button. When you rent intelligence from someone else's platform, their legal exposure becomes your data-retention policy. Unfortunately, there is no setting to turn this off.

Haven't companies always had to protect their secrets?

Yes, and that is exactly why this should feel familiar rather than novel. The discipline of knowing exactly where sensitive information sits, who can see it, and where it must not travel is not new work.

I have spent a career holding other companies' secrets, first in chemical engineering at Air Products, where a product recipe can be the entire business, and later at FreeMarkets, whose model ran on suppliers' and buyers' most confidential volumes and pricing. Protecting information you do not own has been my work for thirty years.

What is new is how the information now walks out the door. Today it takes a copy, a paste, and a free account someone opened last spring. The instinct that protected the crown jewels for decades is the right instinct. It has to be pointed at a new doorway.

The chatbot in the browser tab is a window, not a vault. What you type into it does not stay on your side of the glass unless you put it there.

What should a small or midsize business do first?

Start with the two decisions similar to those you make when you hire someone new: what activities they are expected to perform and what they are trusted to touch. Minimum governance starts with: (1) naming which AI tools are approved for company data, and making them business tiers that do not train on your input by default rather than free personal accounts, (2) writing an acceptable-use policy everyone understands, covering what never goes into a public tool, and (3) naming a person who owns the call when a new question comes up. Settle those three and you have started closing most of the everyday exposure the statistics above describe.

That handles your own people. It does not settle what the vendor may do with what you send, which is a separate decision.

Can an AI company still learn from your data even on a business tier?

Yes, because training is only one of the things a vendor is permitted to do with your data, and the most overlooked one has no opt-out. OpenAI's data processing terms, for example, reserve the right to keep using information derived from customer data once it has been anonymized and aggregated, in order to improve its systems and services.5 Set that beside the training promise and a gap appears: your opt-out protects your identifiable content, but the moment your data is pooled and stripped of identifiers, the terms stop treating it as yours, and the protection stops with it.

Settings will not fix this. You cannot switch off de-identified use the way you turn off training. To prevent it, you have to remove the vendor's ability to derive it in the first place, which is a question of where your data lives.

How do you stop a vendor from using your data at all?

You take the data out of the vendor's reach, and there are four ways to do it, running from a simple request to owning the whole stack.

First, ask for zero data retention. If the vendor stores nothing, there is no copy to aggregate, de-identify, or review. Major providers offer this, though it is gated behind a qualifying use case and approval and can exclude some features, so it is a term you need to request by name.

Second, run the model through your own cloud tenant. Deployed through a service like Azure OpenAI, Amazon Bedrock, or Google Vertex, your prompts stay inside your own cloud environment and, in standard configurations, are not shared with the company that made the model or used to train it. That keeps the model maker, and its de-identified-data clause, out of the picture. Two trade-offs emerge. The cloud provider itself still retains your inputs for a period by default, commonly around 30 days for abuse monitoring, and switching that off means applying for a zero-retention arrangement, which is an approval rather than a setting.6 And a few of the newest frontier models require retention as a condition of access, in some cases sharing the data back to the model maker, so this is a per-model check rather than a blanket guarantee.7 In other words, a condition of using some of the newest models is allowing the model maker to retain your data.

Third, negotiate the contract. A customer with enough size and spend can strike or narrow the de-identified-data right in the agreement itself. That is a real option for a large buyer but rarely realistic for a small business.

Fourth, run an open model on infrastructure you own. When the model runs on your hardware, the vendor never sees your data, so there is nothing to retain, aggregate, or review. This is the only complete answer, and it is the one behind what we call the Intelligence Estate, the case for building AI capability you own, running capable models on infrastructure you control. It is also where the second option leads. Push your privacy bar up to true zero retention and you are often forced off the frontier proprietary models anyway, because some of them require retention to run, which leaves you on open-model-class capability. At that point, running those open models yourself gives you the same class of model with more control and no per-token dependency. For a company between a few million and a few billion in revenue this is no longer exotic, and the hardware to do it now fits on a shelf.

These four are not a ladder you have to climb to the top. Most companies will mix them, keeping frontier models in a controlled tenant for general work and moving their most sensitive workloads onto hardware they own. What each one has in common is that it decides where your data lives, and that is the decision that actually governs your privacy.

Proxigee Services is an AI advisory firm in the Pittsburgh area that helps small and midsize companies take their first strategic steps into AI. That includes making sure the data your business runs on stays the data your business owns. The goal is not to slow anyone down. It is to let your team move fast on tools that keep your secrets in your control.

Your people are going to use AI. The only real choice is whether they use it on your terms or on someone else's.

Want your team fast on AI without giving away the business?

Proxigee Services helps small and midsize companies pick the right tools and tiers, set the guardrails that keep control, and decide what belongs on infrastructure you own.

Schedule a conversation →

Sources referenced

  1. OpenAI, "Enterprise privacy", on API, ChatGPT Enterprise, Team, and Business data not being used to train models by default.
  2. WitnessAI, "AI Data Retention: ChatGPT, Claude, Gemini Policies", updated September 8, 2026, on the training-versus-retention distinction across the major vendors' business tiers.
  3. Cyberhaven, "2026 AI Adoption and Risk Report", February 11, 2026, on the share of AI interactions involving sensitive data and the prevalence of personal-account usage.
  4. The New York Times Company v. OpenAI, preservation order (S.D.N.Y., May 2025; affirmed June 2025; scope narrowed September 2025). Summary: National Law Review, "Privacy Under Pressure: What NYT v. OpenAI Teaches Us About Data Governance".
  5. OpenAI, "Data Processing Addendum", on the reserved right to continue processing de-identified, anonymized, and aggregated data to improve OpenAI's systems and services.
  6. Microsoft, "Data, privacy, and security for Azure OpenAI", on tenant-scoped data, the default up-to-30-day abuse-monitoring retention, and the application-based Modified Abuse Monitoring and zero-data-retention exceptions.
  7. Amazon Web Services, "Data retention for Amazon Bedrock", on retention modes, models that require retention as a condition of access, and the provider_data_share mode.

Terry D. Lyon is the founder of Proxigee Services, a Pittsburgh-based advisory firm helping small and midsize companies navigate the current AI innovation cycle, with a focus on owning your intelligence rather than renting it.

FAQ

Frequently asked questions