People ask us which companies license AI training data, and the lists that answer it often mix businesses that do different things. This map sorts the companies into categories by what each one does, and describes each company using its own published pages, without ranking one company above another.
What kinds of companies are in the AI training data licensing market?
This map groups AI training data licensing companies into seven categories: data marketplaces and aggregators, single-medium licensing programs, collective licenses, publisher access and grounding programs, attribution-based revenue sharing, data collection and annotation providers, and registries that combine opt-out and licensing. Several companies work in more than one category.
The categories matter more than the company names, because a developer looking for a cleared book corpus, a publisher deciding whether to charge AI crawlers, and a creator looking for a way to say no are asking three different questions, and each category answers a different one.
What is a training data marketplace or aggregator?
A training data marketplace or aggregator gathers datasets or content from many owners and makes them available to AI developers under license, either as ready-made datasets or as access to a catalog. Companies in this category differ in the content types they carry, how they source content, and how payment returns to the original owners.
| Company | What it says it does (summarized from its own site) | Source |
|---|---|---|
| Datarade | Describes itself as an AI data marketplace and exchange where buyers find and compare data products from data providers, with AI training data as one of its categories. | https://datarade.ai/ |
| Defined.ai | Describes itself as an AI data marketplace where enterprises buy or commission training data, including speech, text, image, and multimodal datasets. | https://defined.ai/ |
| Human Native AI | Founded to help creators gain control, compensation, and credit when AI uses their work. Its site states that Human Native is joining Cloudflare. | https://www.humannative.ai/ |
| Protege | Describes itself as a source of AI-ready, real-world data for AI development, serving model builders and data holders, with healthcare and medical imaging, video, audio and speech, and spatial data among its categories. Calliope Networks, which aggregates film, television, and news content for licensing to generative AI companies, states on its own site that it is now part of Protege. | https://www.withprotege.ai/ ; https://calliopenetworks.ai/ |
| Troveo | Describes itself as a network of real-world data for AI that licenses non-public data to frontier AI developers, covering video, audio, gaming, text, robotics, and business data, and invites data owners to license their own data through it. | https://www.troveo.ai/ |
| Wirestock | Describes a network of creative professionals that supplies image and video datasets to AI companies, with creators earning commissions when their work is licensed. | https://wirestock.io/ |
What is a single-medium licensing program?
A single-medium licensing program focuses on one kind of work, such as books, or licenses one company's existing library for model training. Because the catalog is narrower, these programs can describe their rights and contributors in more detail, and some give contributors a way to keep their work out of future datasets.
| Company | What it says it does (summarized from its own site) | Source |
|---|---|---|
| Created by Humans | Describes itself as an AI rights licensing platform for books that connects authors, agents, and publishers with AI developers, with rights holders deciding what they are comfortable licensing. | https://www.createdbyhumans.ai/ |
| Shutterstock | Licenses datasets of images, videos, 3D models, and music from its library for machine learning training, pays contributors through its Contributor Fund, and offers contributors an opt-out from future datasets. | https://submit.shutterstock.com/help/en/articles/10594694-shutterstock-data-licensing-and-the-contributor-fund |
What is a collective AI training license?
A collective AI training license offers one set of terms covering works from many rights holders, so a developer can obtain rights without negotiating title by title. Collective programs come from established rights organizations and from newer bodies built for online publishers, and a license covers only the works whose rights holders participate.
| Company | What it says it does (summarized from its own site) | Source |
|---|---|---|
| Copyright Clearance Center | Offers an AI Systems Training License that gives organizations rights to use lawfully acquired works as inputs for training AI systems for external uses. | https://www.copyright.com/solutions-ai-systems-training-license/ |
| RSL Collective | Describes itself as a nonprofit rights organization and licensing platform that uses the RSL Standard to bring collective licensing to online publishers, so they can receive royalties from AI companies. | https://rslcollective.org/ |
What are publisher access and grounding programs?
Publisher access and grounding programs price and control how AI crawlers and AI products reach published content, through per-crawl fees, licensed access for AI answers, or both. Their terms usually attach to access and use at the moment content is retrieved, which is a different transaction from licensing a fixed training dataset.
| Company | What it says it does (summarized from its own site) | Source |
|---|---|---|
| Cloudflare (pay per crawl) | Describes pay per crawl as a private beta feature of AI Crawl Control that lets publishers allow, deny, or charge AI crawlers, while crawler owners can see the price and choose to pay or walk away. | https://www.cloudflare.com/paypercrawl-signup/ |
| Microsoft (Publisher Content Marketplace) | Announced on February 3, 2026: publishers define licensing and usage terms, and AI builders discover and license content for specific grounding scenarios. | https://about.ads.microsoft.com/en/blog/post/february-2026/building-toward-a-sustainable-content-economy-for-the-agentic-web |
| TollBit | Offers publishers tools to monitor, control, and monetize AI access to their content, with a network that also serves AI companies. | https://tollbit.com/ |
What is attribution-based revenue sharing for AI answers?
Attribution-based revenue sharing identifies which sources contributed to an AI-generated answer and pays the owners of those sources a share of the revenue the answer earns. It concerns AI search and answer products rather than training datasets, and payments depend on the attribution method each provider uses to credit sources.
| Company | What it says it does (summarized from its own site) | Source |
|---|---|---|
| ProRata | Builds attribution solutions for AI search and shares revenue with the publishers and creators it works with. | https://prorata.ai/ |
Is data annotation the same as data licensing?
Data collection and annotation are a different business from licensing existing works. Annotation providers label, collect, or produce data to a developer's specification and often supply human evaluation of model output. Some of these companies also sell licensed datasets, so a buyer should ask which rights come with each dataset.
| Company | What it says it does (summarized from its own site) | Source |
|---|---|---|
| Scale AI | Describes its work as spanning the AI stack, from the data that trains models to evaluation and deployment, with humans in the loop. | https://scale.com/ |
| Shaip | Offers data collection in audio, video, image, and text formats, data annotation by domain experts, and an off-the-shelf catalog of datasets available to license. | https://www.shaip.com/ |
Where does Credtent fit?
Credtent is a registry and licensing agent that works with both sides of the market. Creators and rights holders register works for a free opt-out, licensing, or both, and approve every deal before it executes. AI companies license named collections with a manifest and a disclosure pack. Credtent is a Delaware Public Benefit Corporation and builds no AI models.
| Company | What it says it does (from its own site) | Source |
|---|---|---|
| Credtent | Runs the Independent Creative Registry, where creators and rights holders register works for opt-out, licensing, or both; delivers registered opt-out notices to AI companies with a dated record; licenses named collections to AI companies with a manifest and a disclosure pack; and, as of October 2026, routes 85% of gross licensing revenue to rights holders. It also certifies works and organizational processes through its Creative Origin badge system, which signals the role of AI in a specific creative work. | https://credtent.org/about.html ; https://credtent.org/opt-out ; https://credtent.org/for-ai-companies.html ; https://credtent.org/for-content-owners.html ; https://credtent.org/badges/ |
Because we publish this map and appear on it, we held our own entry to the same rule as everyone else's: it says only what our published pages say, and it sits in a category of its own because our published pages describe a registry, a licensing agent, and work-level certification together, which does not fit any single category above. If a listed company shows us a better fit, the map changes.
How was this map built?
Each description summarizes what the company says about itself on its own website, read on October 5, 2026, with the source page listed beside it. Companies appear alphabetically within each category. Inclusion is not a ranking or an endorsement, no company paid to appear, and we correct any description a listed company shows us is out of date.
To request a correction or an addition, write to hello@credtent.org with a link to the page we should read.
What should an AI company ask a licensing provider?
Ask which works the license covers by name, who warranted the rights to each work, whether it covers training, grounding, or both, what records arrive for EU AI Act and California AB 2013 filings, how opt-outs are honored, and how payment reaches the original creators. Those answers tell you more than the category a provider sits in.
Our guide to how AI training data licensing works covers each of these questions in more depth, including what the EU AI Act and California's AB 2013 ask developers to document.
What should a rights holder ask before listing with a provider?
Ask whether you keep ownership, whether the arrangement is exclusive, whether you approve each deal, which uses and buyers you can exclude, how revenue is split and reported, how often you are paid, and how you can leave. The provider's contract answers each of these questions, so read it before you sign.
We would rather you read ours than take our word for it. Our terms for rights holders are summarized on For Content Owners and Pricing, and the free opt-out is there for anyone whose answer is no.