How AI Companies Can Pay Fair Rates for the Content They Need
Summary.
Frontier AI models are trained on the accumulated digital output of humanity, which AI companies acquired essentially for free. This is a problem for both AI companies and content creators. For AI companies, future training will require new, high-qualityThe fight over the data that trains artificial intelligence has become one of the defining economic conflicts of the decade. Publishers, authors, and visual artists argue that their work was taken without permission or payment. AI companies counter that training on available data constitutes fair use and that even if a market in data were desirable, compensating millions of creators is technically impossible: the cost of figuring out what any given piece of data is worth, researchers have argued, would swallow most of the value that data creates in the first place.
Both sides stand to benefit from a fair solution to this impasse and the creation of a sustainable market for content. Any resolution must take both positions seriously while seeing past their literal inconsistency.
On the compensation issue, while content creators are justified in defending their livelihoods, creating a market that ensures fair compensation going forward will arguably serve them better than being paid out for past infractions as existing lawsuits have focused on. And for AI companies, a high-quality continued supply of data that the sector needs for future models, together with legal certainty, is worth more than whatever they save by not paying creators now. As to the technical feasibility, while the data-valuation techniques proposed in research so far are impractical, industry leaders have known since at least 2021—per documents from Anthropic’s Chris Olah and Dario Amodei that surfaced in legal discovery for one of the lawsuits against the company—that low-cost methods exist that could create a thriving market.
This article, which is based on our research, describes how a sustainable market for compensating content creators could work, and why it addresses an important part of the social and economic concerns about an AI future. The fact is that AI companies already produce the two data sets required for pricing content, as a matter of course, every time a model is trained.
The first is the data mixture. This is the proportions in which a model builder blends different kinds of data, which reveal the relative value of each source. For example, quality journalism may be weighed more highly than comments on social media, indicating that it is more valuable—and putting a specific number on how much more valuable. The second is scaling laws. These are the empirical regularities that AI researchers estimate to predict how model performance will respond to additional data and compute. Such estimates, together with economic theory, reveal what share of a model’s total value can be attributed to its training data. Together they show how to slice the pie and how big it is.
These figures make payments to content creators reasonably straightforward. Existing creative industries have successfully used collective management organizations (CMO) to administer payments to individual contributors. There’s no reason not to believe that they can’t be adapted into workable infrastructure for a critical new data market.
Here’s how it would work—and why it would benefit both content creators, AI companies, and the broader public.
What’s at Stake?
Today, everyone is losing. Frontier models are trained on the accumulated digital output of humanity, which AI companies acquired essentially for free. But that stock is running down, and future training will require new, high-quality human data. Research on “model collapse” shows that when models train on the output of other models, quality degrades—sometimes sharply—as outputs homogenize and lose nuance, detail, and unusual cases. As synthetic content increasingly floods the open internet, the expansive scraping that AI companies used in earlier training will increasingly eat its own tail if it does not reward valuable content from all sources, including when generated by the models themselves.
The only durable remedy for training new models is a continued flow of fresh, quality inputs: new journalism, research, code, and especially rich multisensory data from physical tasks models have barely begun to understand, as well as from model outputs. That flow depends on economic institutions such as newsrooms, publishers, universities, studios, workplaces, and unions, that give people the means and the reason to produce it. When AI firms pay nothing for this input, they’re only getting a bargain in the short term. In the long-term, they’re drawing down the stock they depends on, like a clear-cutting logger. Seen this way, data compensation is less a tax on AI than an investment in AI’s own continued capability.
On the other side, creative and information businesses currently face a choice between modest one-off licensing deals and copyright litigation. These two paths that, whatever their merits, do little to tie a creator’s fortunes to the technology’s success. Our framework offers something the current options do not: a share of value scales with that success instead of fighting it.
Most importantly, fighting over what has been done up to this point misses the forest for the trees. So far AI models have earned only about $15 billion in operating profits—too little, however allocated, to make a structural difference to any industry. But in the that future industry leaders are projecting, where AI earns tens of trillions annually, compensation to creators could become an engine of a rich creative future. In the scenario where AI significantly reorders the economy, that’s an attractive alternative to the welfare check of universal basic (or even high) income that has been suggested by AI leaders like Sam Altman as a panacea for potential devastation of labor markets. By giving large populations a productive role, agency, and a stake in AI’s future, it would avoid the pathologies of dependence that easily undermine democracy and dignity.
Ask the Bot
Our framework rests on three ideas, each of which applies a standard economic principle to the behavior of the models.
Data mixture weights tell us how to divide the pie.
Before training, builders decide how much of each kind of data to feed the model—web text, books, code, scientific papers, news, and so on. They tune these proportions carefully, because the mixture materially affects quality, and they guard their recipes closely. Here we can apply the “equimarginal principle,” which is one of the oldest results in the economics of production. It states: if a builder has optimized the mixture, then the last token drawn from each source contributes roughly equally to performance. If news articles were pulling more weight per token than web text, the builder would use more of them until the contributions evened out.
In the context of models and training data, that means that the mixing weight reveals a measure of relative value: sources the model treats as more valuable get more weight. The relationship is not perfectly linear in a deep neural network, but it is informative and quantifiable, and crucially, unlike all the methods thus far proposed, it costs nothing extra to calculate, because the weights must be estimated to train the model anyway.
Scaling laws tell us how big the pie is.
Researchers have documented strikingly regular relationships between a model’s performance and its two key inputs: the size of the model (a proxy for required computation) and the volume of training data. From an economist’s perspective, this offers a map from inputs to output. It shows how much value is created by compute and how much value is created by training data.
Our calculations show that, according to standard industry estimates of scaling laws, data account for roughly 40–50% of a model’s pre-training value, before crediting algorithmic innovation. Treat that as an upper bound. In their memo, Amodei and Olah estimated that figure at roughly 20%. Treat that as a lower bound. Given those bounds, a one-third midpoint is a reasonable working number. In any case, narrowing even to the 20-50% range already puts clear bounds. Two features make this attractive: 1) it is built from evidence that already exists rather than new experiments, and 2) it updates automatically as the technology shifts—rising if models grow more data-hungry, falling if algorithms and computation do more of the work.
Operating profit tells us what to take a share of.
If we’re talking about giving content creators a fair share, the final question here is: share of what? Revenue is the simplest base, but it ignores the real cost of running models and would distort prices and penalize open source/weight competitors. It also ignores the costs of other stages such as post-training and deployment (including payments to other data creators that contribute in those stages). Overall equity is too blunt, sweeping in the value of everything a company does, especially for companies like X and Google that produce many things other than models. The natural middle ground is a share of per-model operating profit—profit over the variable cost of serving a given model, including post-training costs.
This is precisely the arrangement Hollywood uses when it grants creative contributors a share of a film’s profits. It ties payment to the specific asset the data helped build, and it shares both the upside and the risk.
The Music Industry Precedent
More than a century ago, music industry solved a structurally similar problem. Many feared the advent of the gramophone would destroy the livelihood of performers. Collecting royalties from every venue and station that plays a song would be hopeless if each songwriter had to do it alone. So the leading composers, lyricists, and singers established CMOs such as American Society of Composers, Authors, and Publishers (ASCAP) and Broadcast Music, Inc. (BMI) to issue blanket licenses, collect payments, and distribute them according to usage data. Similar arrangements exist (though are less central) in Hollywood, such as the Motion Picture Licensing Corporation. The system is imperfect and its formulas are contested, but it has moved billions of dollars a year for decades.
We’re not the first to suggest CMOs as a possible solution. Recent acts of the European parliament and executive orders from the White House have suggested using them to compensate creators for pre-training data. As an economist and a computer scientist our contribution is how this can work technically and create the right incentives, supporting the leadership of policymakers on the institutional framework.
We see this system operating in three steps:
1. Determine the total payment
Set the level of total payment as a percentage of each model’s operating profit, anchored to the scaling-law-implied data share. Many independent groups, such as academic researchers and evaluation bodies like METR, estimate what that share is, which means it’s hard for any single firm to manipulate the figure. That matters, since firms have a natural incentive to understate it. Model makers could also be made to disclose it as part of the newly established pre-release model review.
2. Divide the pie
Set the distribution across data sources using the mixing weights builders report for their models and other data calculated in the course of model training. This is the part where incentives are mostly aligned: builders already want the mixture right because it drives quality and creators want to be paid fairly for how their work creates value for AI companies. The weights are credible signals of relative value and model makers will have an incentive to continue to improve the quality of this signal to better target their training and to report it to ensure a steady flow of quality data. This also ensures that the bounty of AI value accrues not only to organized intellectual property holders, but to anyone creating data critical to the industry’s future.
3. Distribute the payments
The CMO distributes payments to creators and their representatives—publishers, platforms, guilds—much as ASCAP pays songwriters. These payments also double as a signal as to what kinds of data are valuable to produce. While these initially might simply lead to payments, they might eventually instead become literal equity stakes in the models, offering control rights over model behavior as well as rights to cashflow. This arrangement might actually benefit model-makers, by giving data creators ways to continue steering and improving model performance through adding more data at later model stages (post-training and retrieval-augmented generation).
A Sustainable Future
The technical objection that has kept both sides arguing in the dark, that data simply cannot be valued at scale, does not hold up. The training process already reveals more about what data is worth than the current debate admits. We do not need a precise after-the-fact price for every sentence ever written. We need a workable estimate of relative contributions, a workable estimate of the aggregate, and an institution to turn those into payment: all three within reach.
Today, model-makers and content creators are fighting over the paltry $15 billion in operating profits thus-far generated by the models, which would hardly make a dent in the market capitalizations of the former or the incomes of the later. But those market capitalizations are a hint of what content creators stand to gain by embracing their part in the future rather than defending their declining model of the past: once the dust settles on the present round of IPOs, the value of model makers will likely be close to $10 trillion, a significant share of which could go to content creators. This would be a price well worth model-makers paying to remove two of the greatest threats to the hopes of the industry (backlash and a data drought), especially given many of the industry’s leaders are already searching for ways to share the benefits of their bounty more equitably. This could clear the way for the future leaders like Dario Amodei and Sam Altman have envisioned, where by 2030 AI creates tens of trillions of wealth annually and thus, according to our estimates, funds trillions flowing to content creators every year, a far richer future than their current trajectory.
Thus, there is an opportunity for both content and AI industries to solve their greatest existential challenges and embrace a future of shared abundance. Let us hope they have the leadership to embrace that opportunity.
Recommended For You
Readers Also Viewed These Items