Questions about AI training data licensing reach us from both sides of the table: creators and rights holders register their works with us for opt-out, licensing, or both, and AI companies come to us to license named collections of that work. This guide answers the questions each side brings to us, with the sources we relied on listed at the end.
What is AI training data licensing?
AI training data licensing is a written agreement in which the owner of text, images, audio, video, or data grants an AI developer permission to use specific works to train a model, in exchange for payment and on defined terms. The license names the works, the permitted uses, the duration, and the documentation the developer receives.
Every training license has two parties with different worries. The developer wants to know that the people granting permission had the right to grant it, and that the paperwork will hold up when a regulator or a court asks where the training data came from. The rights holder wants to know what the work will be used for, who else can use it, and how payment reaches the people who made it.
In the European Union, the law gives both parties a reason to want the same document, because the text and data mining exception in Article 4 of Directive (EU) 2019/790 applies only where rights holders have not expressly reserved their rights, and for content published online that reservation can be made by machine-readable means. Once a rights holder has reserved those rights, a developer who still wants the work needs permission, and a license is how permission is given.
What rights does a training license grant, and what does it not grant?
A training license grants permission to use named works to train or fine-tune a model, within the scope the contract states. It does not automatically cover inference or grounding, where content is retrieved to answer a user's question, and it does not transfer copyright. Those uses, and any indemnity, need their own written terms.
The difference between training and grounding matters because some programs now license them separately. Microsoft's Publisher Content Marketplace, announced on February 3, 2026, is built for grounding: in Microsoft's description, publishers set licensing and usage terms while AI builders license content for specific grounding scenarios. A training license and a grounding license can cover the same article and still be two different contracts.
A license also conveys only the rights its signatories hold. A book can carry rights held by an author, a publisher, an illustrator, and a translator, and a license signed by one of them does not bind the others. That is why a careful license comes with a manifest, a list of exactly which works are covered, and with records showing who warranted which rights.
Our For AI Companies page states the boundary in these words: "Licenses include the rights the contracting holders warranted, a named manifest, and a disclosure pack your counsel can use for EU AI Act Art. 53 and CA AB 2013 filings. Those statutes still bind the developer. Indemnity, if any, is only as written in the signed license."
Who sells training data licenses?
Training data licenses come from four kinds of sellers: individual creators, publishers and other rights holders who own catalogs, aggregators and marketplaces that assemble content from many owners, and collective or platform-run programs that offer one set of terms to many participants. They differ in scale, documentation, and how payment reaches creators.
Individual creators can license their own work directly or reach AI buyers through an intermediary such as an agent, a marketplace, or a collective program.
Publishers, studios, labels, archives, and data providers can license a catalog directly. Their main questions are which rights they hold for each work in the catalog and how much control they keep after signing.
Aggregators and marketplaces gather content from many owners and sell it as datasets or as access to a catalog. Some take a license from each owner and resell it; others act as an agent and route payment back to the owners.
Collective and platform-run programs offer standard terms to many participants at once. Copyright Clearance Center, for example, offers an AI Systems Training License to organizations training AI systems for external uses, and platform operators such as Microsoft now run marketplaces of their own.
We keep a neutral map of the companies in each category, described from their own published pages: AI Training Data Licensing Companies: Who Does What in 2026.
How are training data licenses priced?
Training data licenses use four common pricing structures: a flat fee for a defined corpus, a per-work or per-unit rate, a revenue share between seller and rights holders, and usage-based fees tied to how often content is accessed. The price depends on exclusivity, quality, documentation, the rights layers involved, and the license term.
We do not publish a rate card for AI companies. They license from Credtent through custom agreements scoped to the content, the license type, and the size of the organization.
On the rights-holder side our structure is a revenue share. As of October 2026, Credtent routes 85% of gross revenue from each licensing deal to the rights holders and data providers whose content was licensed and retains 15% for the platform, legal infrastructure, compliance tooling, and operations, with royalties distributed quarterly.
Usage-based pricing appears in access and grounding programs. Cloudflare's pay per crawl, which its own page describes as a private beta feature of AI Crawl Control, lets publishers charge AI crawlers for access, and crawler owners can see the price and choose to pay or walk away.
What compliance documentation do AI companies need?
Developers of general-purpose AI models in the EU must publish a sufficiently detailed summary of training content and keep a copyright policy that honors rights reservations, under Article 53 of the EU AI Act. California's AB 2013 requires developers to post training-data documentation, including whether datasets were purchased or licensed. A license should supply records for both.
Article 53(1) of Regulation (EU) 2024/1689 requires providers of general-purpose AI models to put in place a policy to comply with EU copyright law, including identifying and complying with rights reservations made under Article 4(3) of Directive (EU) 2019/790, and to publish a sufficiently detailed summary of the content used for training, using a template from the AI Office. Under Article 113, those general-purpose model obligations apply from 2 August 2025.
California's AB 2013 (Chapter 817, approved September 28, 2024) applies to generative AI systems released on or after January 1, 2022 and made available to Californians. On or before January 1, 2026, and before each later release or substantial modification, the developer must post documentation on its website covering, among other items, the sources or owners of the datasets, whether the datasets include material protected by copyright, trademark, or patent, whether the datasets were purchased or licensed, and when the data was collected and first used.
A developer's filings remain the developer's responsibility. A good license improves the evidence available to the people who write them: a manifest of named works, a record of who warranted which rights, and dates for when the content was delivered. The disclosure pack that comes with a Credtent license is built for that use, and those statutes still bind the developer.
Can a license cover content a model already trained on?
A license can address past training only for the rights holders who sign it. Credtent states the limit in these words: "We can structure a going-forward license plus a covenant not to sue from the rights holders who sign, covering the named works. That is not a release from everyone else, and it is not a regulator cure."
The answer depends on who is at the table. A covenant not to sue is a promise from the signing rights holders, about the named works, and it reaches no further than those signatures. Rights holders who did not sign keep every claim they had, and a regulator's view of a developer's past practices is a separate matter from any private agreement.
How do rights holders license their catalog?
Rights holders license a catalog by documenting ownership, listing the works with a licensing agent or marketplace, setting which uses and buyers they will accept, and approving each deal. With Credtent, owners keep full ownership, approve every deal before it executes, receive 85% of gross licensing revenue, and can choose a free opt-out instead.
Our process for catalogs has four steps. You submit catalog metadata, ownership documentation, and your licensing parameters; we verify ownership and tag the content for discovery by AI buyers; we bring licensing terms to you for approval before any contract executes; and royalties are distributed quarterly with usage reports showing which licensees used your catalog. Listing transfers no copyright and no exclusivity, and you can end our representation at any time. Details are on For Content Owners.
Individual creators have a separate set of plans. The opt-out costs nothing and stays free: we record your reservation of rights and deliver the notice to AI companies, with a dated record of each delivery. Opt-Out Premium, at $49 per year, adds access to the auditing system and does not include licensing. Creator Licensing is $99 per year, and Rights-Holder Licensing, for publishers and others who represent work by many creators, starts at $299 per year. Full terms are on the pricing page.
Opting out and licensing are not permanent choices. A work registered for opt-out can be made licensable later, and a licensable work can be withdrawn to opt-out, with no retroactive fees.
If you are on the other side of the table, For AI Companies explains how a licensing conversation with us starts.