Guides
What is your company's data worth to AI labs?
AI labs now pay for private business records because the public web no longer teaches models what they need. What your data is worth depends on a few clear factors. This guide explains them, shows what public deals tell us, and points you to an estimate for your own company.
Why labs pay for company data
Large language models were trained mostly on public text: websites, forums, books and code. Most of that has been used. Models can write a decent email, but given a real job from start to finish, such as closing a month's books or resolving a support ticket in a CRM, they still fail often.
The reason is simple. Models have seen very few records of how professionals do that work. Your company keeps those records every day, in chat, email, documents, tickets and systems. That is what labs pay for.
What public deals tell us
Few prices for training data are public, and most public deals involve large platforms. They still give useful reference points. The amounts below are in US dollars, as reported.
- Google and Reddit: reported at about US$60M per year for access to Reddit content.
- News Corp and OpenAI: reported at up to US$250M over five years.
- Startups selling old Slack and email archives: reportedly US$10K to US$100K per deal (Forbes, April 2026).
- Reinforcement learning training tasks built from real work: US$200 to US$2,000 each (Epoch AI, January 2026).
- Exclusive deals: reportedly several times the price of non-exclusive ones.
How to read those numbers
A platform with millions of users is a different seller from a 40-person firm. Still, the numbers show two things. First, private work records from ordinary companies already sell, and the amounts reach the tens of thousands. Second, the price per item rises sharply when the data shows complete, checkable tasks, because labs can turn those into training exercises for AI agents.
Seven things that drive the price
Every dataset is valued on the same factors. Some you cannot change, such as how long your company has existed. Others you can influence before you sell.
- Volume and years of history. More records over more years give labs more variety and more examples of the same task done well.
- Structure. Tickets with a clear request, steps and outcome are worth more than loose files. Records that link to each other, such as a ticket, the email thread and the final document, are worth more again.
- Specialisation. Work that few companies do, or that requires expert judgement, is rare in training data. Rare data commands a higher price.
- Language. English data is plentiful. Dutch, Polish, Swedish, Czech and other European languages are scarce, so the same records in those languages are often worth more.
- Cleanliness. Data with few duplicates, little noise and personal details that are easy to remove costs less to prepare, and the seller benefits.
- Exclusivity. If one lab gets the data and no one else does, it pays more. Non-exclusive licenses pay less per buyer but can be sold more than once.
- Number of buyers. Data that several labs want can be licensed several times on non-exclusive terms.
Which sources tend to be worth most
Not every tool in your company holds the same value. Sources that show work from request to result usually rank highest, because labs can turn them into tasks with a known answer.
Support and service desk tickets, project boards linked to documents, and code repositories with their history are often the strongest sources. Internal chat and email come next when they show decisions and problem solving. Procedures, checklists and playbooks add context that raises the value of everything else. Folders of final documents with no history are worth the least on their own, although they still count.
- Usually high: tickets with outcomes, code with history, spreadsheets with working formulas, review comments.
- Usually medium: internal chat and email, CRM records, project boards.
- Usually lower on their own: final reports, marketing material, shared drives with little structure.
Two companies compared
Take two firms of 40 people. The first is a general office that stores mostly final documents in a shared drive, in English, with three years of history. The second is a specialised firm with twelve years of tickets, a written procedure for each type of task, and records in Dutch and German.
The second firm will usually receive a clearly higher offer. Its records are older, linked, specialised and in languages labs have little of. The first firm can still sell, and it can raise its value by adding sources with more history, such as email or a ticket tool.
What lowers the price
Some things reduce value or rule data out. Gaps of several years, scanned documents without readable text, archives full of newsletters and automatic notifications, and text that was itself written by AI are all worth less. So is data where a large share is personal information that has to be removed.
Unclear rights lower the price most of all. If you cannot show that the data is yours to license, a lab cannot use it, whatever its quality. Data you process for clients, content under client NDAs and material written by third parties are left out for this reason.
Why the estimate uses headcount, age and revenue
Before anyone has looked at your sources, three facts predict value fairly well. Headcount says how many people produce records. Years in business say how much history exists. Revenue says something about the scale and complexity of the work.
Lodex's estimate uses these three inputs to give a first range. Across companies, estimates fall within €10K–€600K. The final offer follows from the sources you include, their volume and quality, the access terms and our review.
How to raise the value before you sell
A few practical steps make a clear difference, and most cost little time.
- Choose sources with long, unbroken history over sources that were recently migrated.
- Include the records around an outcome: the request, the discussion and the result.
- Write down your procedures. Checklists, playbooks and review notes add context labs pay for.
- Sort out rights early by checking client contracts and NDAs.
- Leave out noise such as newsletters, automatic alerts and personal folders.
- Mention every language your team works in. Non-English records are often the most valuable part.
How payment works
You agree the scope and price with Lodex in writing before we connect. We connect read-only, clean and package the data, and run a quality review. Once your data passes review, you are paid within 7 days. Buyers see redacted samples only until the package is paid for.
To see a first range for your company, fill in the estimate. It takes about a minute and is not a commitment.
Questions
Can a small company earn anything?
Yes. We work with companies from about 10 people. A small firm with long, specialised history can be worth more than a larger company with little structure.
Is my data sold once or many times?
That depends on the license. An exclusive license goes to one buyer at a higher price. A non-exclusive license can go to several labs. The agreement states which applies.
Why do non-English records pay more?
Models have far less training data in languages such as Dutch, Polish or Swedish. Records in those languages fill a real gap, so labs value them more.
Does the estimate commit me to anything?
No. The estimate is a first range. Nothing is shared until you sign the terms and approve each source.
When is the final price set?
After a call in which we review your sources together. The written agreement states the price, the scope and how payment works.