Guides
What AI labs buy from companies, and why
AI labs now train models to do real work, from filling in an expense report to resolving a support ticket. For that they need records of real work, which mostly sit inside companies. Here is what they buy, sector by sector, and what they turn down.
Why labs buy private data now
The first large language models learned from the public web. That source is largely used up, and it teaches little about how work is done inside a company. In 2026 labs focus on AI agents: models that open a spreadsheet, navigate a CRM, draft a deck and check their own work. To train and test those agents, labs need real examples of the tasks, the steps people took and the result.
That changes what counts as valuable. A finished report is useful. The brief, the drafts, the comments and the final version together are far more useful, because they show the path.
Enterprise workflows and agent training data
The fastest-growing category is data that shows how office work gets done in business software. Labs turn these records into tasks for agents, with a known correct outcome to check against.
- Expense reports and invoice processing, from receipt to approval.
- Spreadsheet work: pivot tables, reconciliations, budgets and forecasts with formulas intact.
- Slide decks built from a brief, with the brief and the drafts.
- Navigation in CRM and ERP systems: creating an order, updating a deal, closing a case.
- Support tickets resolved from first message to solution, including escalations.
Private code with its history
Public code is widely available. What labs lack is private code with the context around it: the ticket that asked for a change, the design document, the commits, the review comments and the tests. Together these let a lab build realistic coding tasks with a checkable answer. Legacy systems, migrations and less common languages are especially wanted.
Expert reasoning
Labs want to capture how professionals think through a problem. How an accountant closes the books and decides what to accrue. How a lawyer reviews a contract and chooses which clauses to negotiate. How an underwriter weighs a credit file. How a planner reroutes a delayed shipment.
This reasoning is often visible in everyday records: review comments, internal email threads, decision notes and the corrections a senior makes to a junior's work.
Labs value the doubts as much as the answers. A thread in which two colleagues disagree about how to book a transaction and settle it shows a model how experts weigh options.
Procedures, playbooks and knowledge bases
Standard operating procedures, playbooks, onboarding guides, quality checklists and internal knowledge bases tell a model how a task should be done and how to check it. Labs use them both to train models and to build tests. Companies often underestimate this material because it feels ordinary. To a lab it is a written description of good work.
Data in smaller languages
Most training data is English. Models are weaker in Dutch, Polish, Swedish, Czech, Finnish and many other European languages, especially for business tasks. Work records in those languages are scarce, so labs pay more for them. A company that works in two or three languages often has an advantage it does not know about.
Local rules add to the value. A Polish payroll procedure, a Swedish employment contract template or a Dutch VAT query thread combines language with knowledge of national law and practice. That combination is very hard for a lab to find anywhere else.
What labs buy, by sector
Each sector has its own valuable records. A few examples:
- Accounting: month-end close checklists, reconciliation threads, Excel workpapers, review notes and client query tickets.
- Legal: clause libraries with fallback positions, review playbooks, matter workflows and internal research memos.
- IT services: private repositories with history, linked tickets, design documents, postmortems and runbooks.
- Consulting: briefs paired with final decks, analysis models, frameworks and proposal libraries.
- Financial services: credit procedures, financial models, compliance workflows and reporting templates.
- Logistics: exception handling threads, TMS, WMS and ERP workflows, customs procedures and warehouse SOPs.
How labs use what they buy
Labs use business data in three main ways. The first is training: the model reads many examples of a task and learns the patterns. The second is evaluation: labs keep a set of real tasks aside to test whether a new model can do the work. The third is reinforcement learning, in which a model attempts a task in a simulated environment and is rewarded when the result matches what a professional did.
Each use favours different data. Training benefits from volume. Evaluation needs well-documented tasks with a clear right answer. Reinforcement learning needs both the task and a way to check the outcome, which is why records with a visible result, such as a closed ticket or an approved report, are in high demand.
How labs judge quality
Before buying, a lab looks at redacted samples. It checks whether the records are complete, whether the steps can be followed, how much noise there is and how consistent the format is across years. It also asks where the data came from and whether the seller had the right to license it.
Good provenance has become part of quality. Under the EU AI Act, providers of general-purpose AI models must publish a summary of their training data. Labs therefore prefer data that comes with a clear, written record of its source, scope and de-identification.
What labs do not want
Labs are selective. Some data has no value to them, and some they cannot legally use.
- Personal data: names, contact details, health data and ID numbers. These are removed or excluded.
- Data with unclear rights: client-owned files, content under NDA and third-party material.
- Noise: newsletters, automatic notifications, calendar invites and duplicates.
- Text written by AI tools, which teaches models little.
- Scanned documents without readable text.
- Records with no outcome, such as a ticket that was never closed.
- Content that is already public, which labs have in most cases.
How to tell whether your data fits
Ask four questions. Do our records show a task from start to finish? Do they cover several years? Do they show expert judgement? Do we own them? If the answer to most of these is yes, your data is likely to interest labs.
Lodex helps you check this before you commit. The estimate gives a first range for your company, within €10K–€600K, and a short call shows which of your sources labs are most likely to buy.
Questions
Do labs buy data from ordinary companies or only from large platforms?
Both. Large platforms sell volume. Ordinary companies sell records of real work that platforms do not have, which is what agent training needs.
Is email and chat history useful?
Yes, when it shows work being done: decisions, problem solving and handovers. Newsletters, notifications and social messages are removed.
Will a lab use my data to build a product that competes with me?
Labs train general models on data from many companies. Lodex licenses only to AI labs, never to your competitors, and nothing is published.
Do labs want data in my language?
Usually yes. Business data in European languages other than English is scarce, and labs pay more for it.
What if most of our records contain customer details?
Customer details are removed during processing. What remains, the steps, the reasoning and the outcome, is what labs pay for.