The aircounsel.ai blog

Before you sell your company's data to an AI firm, answer these three questions

A founder came to us convinced his old startup's data was a goldmine an AI company would pay handsomely for. It might have been. The problem was never the price. It was whether he could lawfully sell it at all — and the reasons he couldn't are ones almost every founder gets wrong.

Back to blog
A stack of data records secured by a padlock — who owns the data you want to sell?

The setup

A founder we advised had built a software business that served other companies. The business wound down. Investors were repaid, the co-founders wrote the remaining assets off as worthless, and everyone moved on. Sitting quietly in a cloud account was a large body of data the software had generated over years of operation.

Then the AI boom arrived. That kind of data — real, high-volume, generated by actual business activity — is exactly what companies training AI models want. Buyers came calling. The founder was ready to license it to several of them through a data broker, on a fee tied to the deal closing. On paper, a tidy way to turn a dead asset into cash.

We had to tell him to stop. Not because the data lacked value, but because three questions had never been asked. Any founder thinking about selling data to an AI company should ask them first.

Question 1: Do you actually own it?

Writing an asset off as worthless is an accounting decision. It is not a transfer of ownership. When a company is wound down, its assets do not automatically become the personal property of a founder — depending on the jurisdiction they pass to shareholders, or even to the state. Holding the data in your personal cloud account does not make it yours.

And if it was called worthless when investors were repaid, and it turns out to be valuable, the people who were told it had no value — co-founders, investors — may have a claim on the upside. The moment money changes hands, that becomes a live dispute. So the first question is not "what is it worth" but "whose is it, on paper?"

Question 2: Whose data is it under the contracts?

This is the one that ends most of these deals. If your company provided software to customers, the data that flowed through that software was almost certainly your customers' data, not yours. Standard software and data-processing terms say exactly this: the customer owns the data, and the vendor is a processor allowed to use it only to run the service. That permission usually ends when the contract ends, and often comes with an obligation to delete the data, not keep it.

So the data you think you are about to sell may belong to your former customers, and you may have been contractually required to erase it. Before anything else, the old customer contracts have to be read. If they say what these contracts usually say, there is nothing to sell.

Question 3: Can it lawfully be sold given privacy law?

Even with clean ownership and clean contracts, data about identifiable people carries a third layer. Personal data cannot simply be sold because someone wants it. The people it describes gave it for one purpose, under one set of expectations. Selling it to an AI company for a completely different purpose usually is not covered by that original consent, and privacy regimes on both sides of the world — from Europe and the US to India's own data protection law — treat selling personal data as a regulated act with real penalties.

Some categories are far worse than others. Recordings of voices, financial information, health data, anything touching regulated industries — these turn a licensing deal into a liability the buyer's lawyers will find, and the seller will wear.

"We'll just anonymise it" is not the escape hatch

The instinct, once these problems surface, is to say: we will anonymise it, or turn it into synthetic data, or train a model on it and sell that instead. It feels like a clean workaround. It is not.

Anonymising solves only the privacy layer, and only if done to a genuine legal standard — which is much harder than stripping out names. It does nothing about the first two problems. If the data is not yours, processing it into a new form does not make it yours. If your contracts barred using it, then building a synthetic dataset or a model from it is itself a prohibited use — and can carry its own separate exposure. A derivative inherits the defects of the thing it was derived from. You cannot launder a rights problem through a clever transformation.

What good advice actually looks like

The honest engagement here was not "help me sell my data." It was "let's find out whether there is anything you can lawfully sell, before you put it in front of a single buyer." That means establishing ownership, reading the old contracts, and testing the privacy position first. Only if those clear does the licensing conversation begin.

It is less exciting than a quick payday. It is also the difference between monetising an asset and selling a lawsuit. The founders who get this right treat the legal groundwork as phase one of the deal, not an afterthought once a buyer is at the table.

If you are sitting on data an AI company wants to buy, that may be genuinely valuable — and precisely because it is valuable, it is worth making sure you own it and can lawfully sell it before you start. That first, unglamorous step is exactly the kind of work we do at aircounsel.

Newsletter

Get contract tips and startup legal insights in your inbox.

No spam. One email per week. Unsubscribe anytime.

By subscribing you agree to receive emails from aircounsel.ai. No spam, ever.