What does it mean to sell anonymised company data?
It means a buyer gets the substance of your work and none of the people in it. An AI lab learns how your team handles a complaint, scopes a project or fixes a bug. It never learns who wrote the message, which client it was about or how to reach anyone.

We buy that kind of data from companies of about 10 people and up: Slack and Teams messages, email, documents, tickets, CRM records, code and procedures. We clean it, then license it to AI labs that train and test models. You keep your originals and keep using them. You license a copy.
Our claim is simple: what we deliver is fully anonymised. That claim only means something if it is backed by concrete steps. This page lists them, and it also says plainly where the GDPR still applies to you.
What is anonymised data under the GDPR?
Anonymised data is information that no longer relates to an identified or identifiable person. The GDPR does not apply to it. That rule sits in Recital 26 of the GDPR, which says the data protection principles should not apply to data rendered anonymous so that the person is no longer identifiable. In US spelling the same idea is called anonymized data.
The test is strict. Recital 26 asks whether anyone could identify a person using all the means reasonably likely to be used, by you or by someone else. It names cost, time and available technology as factors. Removing a name is not enough if the rest of the record still gives the person away.
Anonymisation
Anonymisation is the process of turning personal data into anonymised data, for good. Nobody, including us and including you, can trace a record back to a person afterwards. The result is outside the GDPR.
Pseudonymisation
Pseudonymisation is defined in Article 4(5) of the GDPR. It means processing data so it can no longer be linked to a person without additional information, which is kept separately and protected. Swapping "Anna de Vries" for "Employee 0042" and keeping the list that maps one to the other is pseudonymisation.
Recital 26 is clear that pseudonymised data is still personal data, because it can be linked back with that extra information. That is why we keep no list, no lookup table and no mapping of any kind.
Anonymisation vs pseudonymisation: what is the difference?
The difference is the key. Pseudonymised data has one, somewhere. Anonymised data has none, anywhere. Everything else follows from that.
| What is kept | Can it be linked back? | GDPR status | |
|---|---|---|---|
| Raw data | Everything: names, emails, phone numbers, client names, IBANs, addresses | Yes, it already names people | Personal data |
| Pseudonymised | The text, with names swapped for codes or tokens | Yes, by whoever holds the key | Still personal data (Recital 26, Article 4(5)) |
| Fully anonymised (what we deliver) | The text, with identifiers replaced by type tags such as [NAME] or [CLIENT] | No. No key is kept by anyone and revealing context is removed | Outside the GDPR once no one can reasonably identify a person (Recital 26) |
A practical example: in a pseudonymised file, "Employee 0042" appears in 300 messages, so a reader can follow one person across the whole archive. In our files every name becomes the same tag, [NAME]. The conversation still makes sense, but there is no thread to pull.
Pseudonymised data has a key somewhere. Fully anonymised data has no key anywhere.
What do we remove before data is sold?
We remove every direct identifier, and then the context that could point to one person. The first part is mechanical. The second takes judgement.
- Names of employees, customers, suppliers and contacts, including nicknames and @mentions
- Email addresses, phone numbers and postal addresses
- Customer and supplier company names
- Bank account and IBAN numbers, invoice and contract numbers that lead to an account
- Customer numbers, user IDs, licence plates and other identifying numbers
- Signatures, email footers and profile details
- Context that points to one person, such as "our only tax lawyer in Leeds" or the date of someone's sick leave
Each item is replaced by a tag for its type, so the text keeps its shape. The examples below show what that looks like on everyday records.
- Slack, #sales
Anouk, can you call Pieter Verbeek at Brightwater Foods before Friday? Direct line: +31 6 12345678. They want last year's price.
- Jira ticket
Invoice export fails for Kroonstad Logistics. Reported by Tom de Wit, contact l.jansen@kroonstad.example. Reproduced on staging.
- Email
Hi Marieke, please pay the deposit to NL00 BANK 0123 4567 89 by Monday. Keys are at our Utrecht office. Regards, Joost Bakker
- CRM note
Met Sanne Vos from Hollander Bouw in Zwolle. Wants a quote for 40 seats. Follow up at sanne@hollanderbouw.example.
Every tag is generic. No key back to the original is kept.
Context is the hard part. "Our CFO" names one person in any company. A job title in a team of three, a town plus a rare event, or a quoted salary can do the same. Where a detail adds little for a buyer, we drop it. Where it matters, we make it more general, for example "a senior manager" instead of a title held by one person.
What stays out of the dataset entirely?
Some data is never part of a sale, cleaned or not. The risk is too high or the data is not yours to sell.
- Health data, including sick notes and medical details in HR files
- National ID numbers, passport numbers and similar government identifiers
- Other special categories under the GDPR, such as religion, trade union membership, sexual orientation and biometric data
- Client data you hold as a processor, for example payroll you run for clients or data you host for customers
- Privileged or client-confidential material, unless it can be fully stripped
Professions with a duty of secrecy need extra care. Our page on selling data from a law firm shows what stays in and what stays out. Templates, workflows and internal know-how are usually still valuable without a single client detail.
How do you anonymise Slack, Teams or email data?
We combine automated detection with human review. Software finds the obvious patterns at scale: names, email addresses, phone numbers, IBANs, postcodes. People then read samples of the output and look for what software misses, such as context that points to one person.
Slack and Teams
Chat data carries identity in more places than the message text. Display names, @mentions, user IDs, channel names like #client-brightwater and shared file names all get the same treatment. Emoji and timestamps stay, because they show how work flows, but exact times can be coarsened where they would identify someone. Our pillar on selling Slack, Teams and Jira data explains what labs learn from each tool.
Email hides identifiers in headers, signatures, legal footers and long quoted reply chains. We clean the From, To and Cc lines, strip signatures and footers, and treat every quoted reply as text to be cleaned again. A name that was removed from the top of a thread must not survive in a quoted message at the bottom.
Tickets, CRM and documents
Tickets and CRM notes are full of customer names, account numbers and contact details. Documents add headers, author fields and file metadata. We clean the visible text and the metadata, because a file's author field can name a person just as well as its first line.
- We agree in writing which sources are in and which are out.
- Automated detection tags identifiers across the full set.
- Tagged items are replaced by type tags such as [NAME] or [IBAN].
- Reviewers read samples and look for revealing context.
- Whatever they find is fixed across the whole set, not just in the sample.
Can anonymised data be re-identified?
Badly anonymised data can. Well anonymised data should not be. The Article 29 Working Party, where the EU's data protection authorities worked together, set out three ways it goes wrong in their Opinion 05/2014 on anonymisation techniques: singling out one person, linking records about the same person, and inferring facts about a person from other details.
Our steps map onto those three risks. Generic tags stop linking, because every name looks the same. Removing revealing context stops singling out. Leaving out health data and other special categories limits what anyone could infer.
The key matters most. If a list linking tags to real names existed, anyone who got hold of it could undo the work. So we never create one. Nobody holds it: not us, not you, not the buyer. Once the originals are deleted, there is nothing left to match against.
No method is perfect, and we would rather say so. That is why a person reviews the output, why you check a sample, and why we license only to AI labs and research teams under contract. Nothing is published.
What do you check before anything is sold?
You check a cleaned sample and sign it off. Nothing is sold without that approval. The steps below are the end of how a sale works from handover to payment.

- Secure handover: encrypted, read-only access to the sources you approve, under NDA from day one.
- We clean it: names, customer details and other identifiers are removed during processing.
- You approve: you read a cleaned sample. If you spot something, we fix it across the set and you check again.
- Originals deleted: once processing is done, we delete our copy of your originals. Yours stay where they were.
- Paid: the money is in your account within 7 days after the data passes review.
Buyers see redacted samples only. The full package is released after they pay. We never license to your competitors.
Anonymisation does not wipe out the value. Labs want the reasoning, the back and forth and the decisions, not the names. Company data can fetch €10K–€600K, and larger or specialised datasets can earn up to hundreds of thousands of euros. Our page on why business data is worth money explains what drives the price.
Do employees have to be told?
Yes. The finished dataset is anonymised, but making it is not. To clean your messages, someone has to process them while they still contain names. That step is processing of personal data under the GDPR, and the Opinion 05/2014 above says the same: anonymisation is itself further processing.
Anonymised data is outside the GDPR. Making it is not.
In practice that means three things on your side. You need a documented assessment: a compatibility test for the new purpose, usually a legitimate interest assessment and often a DPIA. You need transparency: an updated privacy notice that tells staff and contacts that de-identified copies may be licensed for AI training, with a simple way to object. And if you have a works council, involve it early.
Our guide on whether it is legal to sell company data under the GDPR walks through each step with a checklist. The scope and terms we sign with you also give you a record for your own files.
Questions
Is anonymised data still personal data under the GDPR?
Not if it is truly anonymous. Recital 26 of the GDPR says the rules do not apply to data where a person is no longer identifiable by any means reasonably likely to be used. Pseudonymised data, where a key still exists, remains personal data.
What is the difference between anonymised and anonymized data?
None. Anonymised is the British spelling and anonymized the American one. Both mean data that can no longer be linked to a person.
Do you keep a key so you can trace data back if needed?
No. We replace identifiers with generic tags such as [NAME] and never create a list that maps tags to people. Keeping such a list would make the data pseudonymised, and so still personal data.
Can I see what my data looks like after cleaning?
Yes. You check a cleaned sample before anything is sold, and nothing goes ahead without your sign-off. If you spot something, we fix it across the whole set.
What happens to my original data?
You keep it and keep using it, because you license a copy. Our copy of the originals is deleted once processing is done.
Is medical or HR data ever included?
Health data, national ID numbers and other special categories stay out entirely. HR records with that kind of detail are excluded, not cleaned and sold.
Does anonymisation make the data worthless to AI labs?
No. Labs train on how work gets done: the questions, the reasoning and the decisions. Names and phone numbers add nothing to that, so removing them costs very little value.
Who buys the anonymised data?
AI labs and research teams that train and test models, under contract. We never license to your competitors and nothing is published.
This is general information, not legal advice.