New: who buys company data, October 2026 list
Data Licensing Report

Question

What does “anonymized” mean, and could it be traced back to us?

We may earn a referral fee when a business we introduce completes a deal. Telegraph Lab is a commercially affiliated provider.

Data Licensing Report may earn a referral fee when a business that applies through this site completes a deal with a participating provider. Telegraph Lab is a commercially affiliated provider: the owner of Data Licensing Report is paid commission on some Telegraph Lab deals, and does not own Telegraph Lab. Providers are listed alphabetically and described from their own public materials using the same fields.

How we make money
On this page
  1. Is “anonymized” the same as “de-identified”?
  2. What do providers say they remove, and who sees our records first?
  3. Could someone figure out it’s us?
  4. Our bids and pricing are our whole edge. Can a competitor see them?
  5. There’s personal stuff in there: divorces, medical leave. What happens to that?
  6. Is this how they get my customer list?
  7. What should I ask a provider before I sign?
  8. How can I see what a redacted copy of our records would show?
Why trust us

“Anonymized” usually means names, contact details and account numbers are removed or replaced. Business records can still point back to your company or a person through context, so check what is removed, who removes it, and what the license forbids.

Is “anonymized” the same as “de-identified”?

No. De-identification is a process; anonymous is a result, and the two are held to different standards. NIST, the federal standards agency, defines de-identification as “any process of removing the association between a set of identifying data and the data subject”, and its 2023 guidance suggests re-identification studies to gauge the risk that remains. Anonymous is a higher bar. EU law treats information as anonymous only when no one can identify the person by “all the means reasonably likely to be used”, and FTC technologists wrote in July 2024 that data “is only anonymous when it can never be associated back to a person”.

Term What it means Who defines it
De-identified Identifiers removed or replaced. Risk is reduced, and the remaining risk is something to measure NIST SP 800-188, September 2023
Pseudonymized Each identifier replaced with the same label every time, such as Person 01. Under EU law this is still personal data, because extra information could tie it back to a person GDPR, Recital 26
Anonymous No one can tie it to a person by any means reasonably likely to be used GDPR, Recital 26; FTC, July 2024
Deidentified data (Texas) Data that “cannot reasonably be linked to an identified or identifiable individual”, or to a device linked to that person Texas Business and Commerce Code §541.001

Most provider pages use “anonymized” and “de-identified” as the same word. Read both as de-identified until a provider shows you more. The consistent labels that keep records useful, so the same person is Person 01 in every thread, are explained in what a data sample is; under the EU definitions they are pseudonymization, not anonymization.

Texas’s privacy act requires a covered business holding de-identified data to “take reasonable measures to ensure that the data cannot be associated with an individual”, publicly commit not to re-identify it, and contractually bind every recipient to the act’s rules (§541.106). The same three parts make a useful check on any provider: what it removes, whether it commits in writing not to re-identify, and what its license forbids recipients to do. The rest of what Texas and federal law ask of a seller is in is it legal to sell company emails.

What do providers say they remove, and who sees our records first?

Most lists cover names, emails, phone numbers, addresses and account numbers. The larger difference is where the work happens, because whoever de-identifies your records sees them whole first. As of October 2026:

Provider What it says it removes Where it happens, and who sees raw records
Appen Personal data, with “synthetic rewrite where workflow structure has to survive the redaction” An automated pipeline; it says “No lab and no third party sees your raw data”
Avelence Leaves de-identification to the partner you select, which explains its own process (Avelence) The selected partner
Corpus Names, emails, phone numbers, addresses, government IDs, payment details and credentials (Corpus) Not stated; you review a redaction report with samples and sign off before transfer (Corpus)
Handshake AI Identifying details, plus cleanup of the records (Handshake AI) Handshake AI; you “transfer the data as it exists today”
Mercor Says its platform “extracts and de-identifies” data; the method is not described Mercor’s platform; “No AI lab buyer or third-party vendor ever sees your raw data”
micro1 Personal information where relevant; some datasets are synthetically rewritten (micro1) micro1, with access restricted to a limited number of authorized staff (micro1)
Polyshares “Names, emails, phone numbers, account numbers, and addresses”; keeps roles, timestamps and the work An unnamed specialist PII partner, before Polyshares or a lab sees the record; it also says only two people at Polyshares see a company’s identity alongside its data (Polyshares)
Replay Says “only the cleaned output is ever licensed” Replay; raw data comes to it for de-identification (Replay)
Scale AI Personal information, so records can’t be traced to individuals or accounts (Scale AI) “inside your environment, before any data is extracted”
Telegraph Lab (affiliated with this site) Names, contact details, account IDs and confidential information, excluded or replaced with consistent placeholders (Telegraph Lab) Telegraph Lab coordinates the export and preparation (Telegraph Lab)
Troveo Names, contact details, credentials and client identities, “under one documented standard” that is not published Troveo, from the exports you run (Troveo)

Two providers state the limit outright. Telegraph Lab’s FAQ says de-identification “reduces risk; it cannot guarantee that re-identification is impossible”, and Avelence says it “does not settle every permission or confidentiality question”. Ask every provider who sees the raw records, how many people, and when that copy is deleted.

Could someone figure out it’s us?

Possibly, from context, unless your company’s own names are on the replacement list. Most published lists target people and contact details. Your company name, email domain, letterhead, truck numbers, project names and the town you work in are context, and someone who knows your market can put them together. NIST’s guidance treats this as a disclosure risk “to individuals and establishments”, meaning businesses as well as people.

For example, a hypothetical thread about re-roofing a named distribution center after a dated hailstorm, with the crew size and the contract value, could identify the contractor to anyone in the local trade even after every personal name is replaced.

Providers make promises about naming you:

Those promises cover labels and listings. What the records themselves reveal depends on what was removed. The sale may not stay private either. Frankfurt Kurnit notes that California’s AB 2013 requires generative AI developers “to publicly disclose high-level information about their training data, including its sources and whether it was purchased or licensed”.

If your company name, domain, main customers and project names are on the replacement list, a reader sees Company 01. If they are not, a reader may see you. Ask for the list in writing and add your own terms to it.

Our bids and pricing are our whole edge. Can a competitor see them?

De-identification removes names and contact details, not prices, so bids, unit prices and margins stay in the copy unless you leave them out or have the amounts replaced. Who may read them is set by the license. Telegraph Lab says the agreement “identifies authorized recipients and permitted uses, including any access by frontier AI labs”; Handshake AI says data is “used for model training and evaluation only, and is not published or made publicly accessible”; Polyshares says “Resale or use beyond that scope is not permitted”.

Some risks sit outside any contract. Frankfurt Kurnit’s October 1, 2026 commentary says that once data trains a model “it cannot practically be pulled back out”, that models “can also reproduce portions of their training data in outputs to other users”, and that a sale may “jeopardize trade secret protection for pricing logic, strategy documents, and source code”.

  • If pricing is your edge, keep estimating files and bid tabs out of scope. Leaving out records explains how to write that into the agreement.
  • If you license them, ask whether the buyer will take them with amounts replaced, and get the named recipients and permitted uses in the agreement.
  • To see the difference, Data Licensing Report’s sample tool keeps money amounts by default, because prices help a buyer judge records; switch that off and each amount becomes [AMOUNT].

There’s personal stuff in there: divorces, medical leave. What happens to that?

The names come out, but the details of what happened stay in. Frankfurt Kurnit notes that de-identification “addresses whether a record can be linked to an individual, not whether its content is confidential”, and a disciplinary thread or pay complaint “remains sensitive without a name attached”. For example, in a hypothetical 60-person company, “Person 07 is out until March after surgery” points to one person for anyone who worked there that winter.

Unions made the same argument in the Spirit Airlines bankruptcy sale. The flight attendants’ union worried that Google’s de-identification might still preserve “referential integrity across the data set”, the links that let records be tied back to people, and asked to “remove all identifying information that can be traced back to individuals and employees”, Fortune reported on August 21, 2026. Two more unions joined the objection on August 28 (court filing), and as of late September 2026 the sale still needed the court’s approval (Moneywise).

Medical leave carries a legal rule. The FMLA covers private employers with “50 or more employees in 20 or more workweeks” in the current or previous calendar year, and requires records about medical certifications and histories created for FMLA purposes to be kept “as confidential medical records in separate files/records from the usual personnel files”. A supervisor’s email asking when someone is back from leave sits outside that separate file, in the mailbox you might license.

What programs and advisors say about this material:

  • SimpleClosure, whose AssetHub began with closing companies, says its process excludes “sensitive employee and health records” and anything the seller marks out of scope.
  • In its Spirit bid, micro1 proposed excluding sensitive employment material, Fortune reported.
  • Frankfurt Kurnit lists HR and personnel files, payroll and tax records, and health and benefits information among the categories to keep out (Frankfurt Kurnit).

If HR, payroll and benefits have their own mailboxes or folders, leave them out whole. If personal matters are scattered through ordinary email, search for the words that flag them (leave, doctor, divorce, garnishment) and pull those threads before anything is exported, then read the sample. Leaving out records covers how to write that into the scope.

Is this how they get my customer list?

A de-identified copy is not supposed to contain one, but check how “customer” is defined. Lists built around personal data remove people and their contact details, and a customer that is a company, such as a GC or a building owner, is not a person. Troveo names client identities on its list (Troveo) and Polyshares says customer names are stripped (Polyshares); for any other provider, ask whether customer company names come out.

Buyers say they want the work itself, not the contacts. Avelence says AI buyers license company data “for the context behind real work”, and Sell My Business Data says the buyer gets business patterns rather than a list of people or customers (Sell My Business Data).

Your customer list is exposed at two other points. The first is before de-identification. Handshake AI and Nyne ask you to transfer data as it exists, and Replay receives raw data to de-identify it, so they handle the names first (Handshake AI, Nyne, Replay). The second is the provider’s other business. Check each provider’s privacy policy for whether it also collects or sells contact data; Nyne’s, updated September 1, 2026, says it is “registered as a data broker with the California Privacy Protection Agency, the Texas Secretary of State, and the Vermont Secretary of State”.

Ask in writing whether your raw or de-identified records can feed any other product, and when the raw copy is deleted. Replay says data is deleted on request if you do not continue (Replay), and micro1’s application page says original datasets are deleted after processing (micro1). In spreadsheet exports from business software, you can leave the customer contact columns out altogether. Why a company would approach you at all is covered in what’s the catch and is it a scam.

What should I ask a provider before I sign?

Ask what is removed, who removes it, what you see before delivery and what the license forbids:

  1. What is removed. The written field list, and whether your company name, email domain, customers’ and GCs’ names, claim or account numbers, and money amounts are on it.
  2. Who removes it, and where. Inside your systems, at the provider or at a third party; who sees the raw records, how many people, and when that copy is deleted.
  3. What you see first. Corpus sends a redaction report with samples for sign-off (Corpus), Appen a “representative sample pack”, and Scale AI asks for “Your sign-off on every data package”. Ask for the equivalent from any program.
  4. What the license forbids. Re-identification, resale, publication and any recipient not named in the agreement. A clause like License My Data’s, under which “re-identification is prohibited by contract”, is what to look for.

How can I see what a redacted copy of our records would show?

Make one yourself first. Data Licensing Report’s free sample tool runs in your browser. It reads a Slack export, a mailbox or spreadsheet exports, replaces names from headers and columns, contact details and your own list of terms (company name, main customers, project code names) with consistent labels, and shows you every item it picked. You can leave out whole channels, mail folders or columns, turn money amounts and dates into placeholders, and use Show originals to compare each replacement on your computer. Nothing is uploaded unless you choose to, and the downloaded package never contains the original values.

Names typed inside messages are the hardest to catch, and free-text detection can miss some, so read each item. Then compare a provider’s de-identified output with your own; whatever it left in that you removed is the next question to ask.

Providers named on this page

Get offers

Find the next step for your company’s data.

  • No records or exports needed to apply
  • Free for businesses
  • Your company profile comes to our team for review

Already have an offer? Compare it

Start your application

Do you keep at least 3 years of email or business-software history?

We may earn a referral fee if a deal closes. How we make money