Arguments in Favor of AI Fair Use Kevin Kelly, founder of Wired magazine, argues in a blog post that training large language models on copyrighted material should be considered fair use under US copyright law, citing the transformative nature of the latent space and the emergence of new capabilities. Kelly contends that LLMs do not store copies of the original works, that the transformation is one-way and irreversible, and that the new uses enabled by LLMs are fundamentally different from reading the source material. Arguments in Favor of AI Fair Use https://kk.org/thetechnium/arguments-in-favor-of-ai-fair-use/ I made some notes on the nature of the training of LLMs, and about whether we as a society should consider the material used to train them as a fair use of that material. I believe it would be best for us to consider them fair use, and I made some points in its favor. The context of these points is US copyright law, which I did not spell out here. There are other commonly used arguments pro, and many arguments against, which I have also not listed. I wrote these points as a way for me to think aloud along my own lines, and they are not arranged in any order to make a tight case. I welcome constructive comments, or points I may have missed. 1. Text, images, or music, etc, are all expressions made by humans, which are copyrightable. During training LLM transforms these expressions into an abstraction called “latent space.” It takes one kind of thing – expressed content – and transforms into another kind of thing – an almost mathematical abstraction that contains no expression. This is why LLMs are called transformers. The latent spaces of an LLM are closer to something like a syntax which is not copyrightable. 2. This latent space transformation is a fundamental transformation, because it goes one-way. It is an asymmetrical process: Content can be moved into latent space, but the latent space cannot be reversed to go back into the original content. It has truly been transformed. 3. The usage patterns of how the general public uses the transformed material is new. When an LLM is asked to troubleshoot a broken electronic device, this service does not resemble reading a book or a newspaper article. When an LLM is used to develop a marketing strategy, it is a profoundly different action than reading text. When an LLM is assigned the task of proving a math problem, it is not substituting for a book about math. The transformations accomplished by LLMs are both in their form and in their use. Not only are they transformed, but they are transformed into a whole new category we have not made before. 4. The key benefit of LLMs is their smartness, which emerges from the latent space and is not present in the source material itself. Their ability to pass a history exam is not present in any history book, and is not even present in ALL history books. These new benefits emerge from the process of creating an LLM trained on history books, and then transformed into a latent space, which produces this smartness. Much, if not most, of the value of LLMs derives from their ability to generate new benefits beyond what the source material alone can yield. 5. The information inside an LLM is not stored in the form of copies of things. Even though an LLM may know the full content of millions of books, it does not contain within it any copies of the books. It may be able to recognize any object in a picture, even though it stores no pictures. Instead of containing a copy of the image, it contains the information contained in the image. The weights of an LLM can be copied, but the latent space itself does not resemble a copy. The latent space is a non-copy entity. 6. Intermediate, transitional copies are the norm in the digital world. When your phone pulls up a web page it is technically making a copy of that page for a brief moment. When you send an email, it is copied in transit by telecom companies many times, but we don’t count them as invoking copyright. Since these intermediate copies are not stored, we don’t constrain them. Courts have already permitted literal copying that never surfaces publicly, The copies that LLM’s read once during training, are not stored, are not surfaced, and are therefore normal transitional copies that are fair use. 7. The genius of the LLMs comes in part from the astounding fact that all bits of information from millions of different kinds of sources and subjects are mapped onto a single “map”. There is a single conceptual space the latent space where every fictional story, and every bit of medical information, is combined with all geological knowledge, and all news events. And so on. As a result the value of any particular contribution to the overall value of the LLM is almost incalculably small. 8. Any individual item used in training the LLM has a very low value because 99% of almost every source is redundant. There is very little unique information in any item, even for seemingly original material. When an additional piece of content is inserted into the conceptual space of all human knowledge, most of it is redundant with other material. That means, removing an average bit of source content would make virtually no difference to the output of the LLM. As LLMs continue to scale up, the value of any particular source continues to diminish. 9. Modern LLMs have trillions of parameters – that is they have trillions of attributes they are describing. The consequence of mapping all human knowledge into one latent space with trillions of attributes is that every bit of information intersects, or influences, every other bit of information. Calculating this enormously complex relationship between trillions of influences is the grand task of a massive data center. This calculation is the most complex computing task we have ever done. But it is so complex that we cannot practically unravel the ripple of influences of any particular source. The degree of influence for each source content is therefore not calculable at this scale. It is not dissimilar to our considerable inability to calculate in a quantifiable way the influences in our own thinking and knowledge. Therefore the import of a particular source cannot be practically measured. 10. In creating its knowledge, an LLM does not distinguish between sources of influence in its training. There are more hours of Star Wars reviews than Star Wars movies. You will have a lot of knowledge about Star Wars even if you have never watched a copy of a single one. What an LLM knows about Star Wars may come primarily from sources other than the movies themselves, from authors who have watched Star Wars. These diverse secondary and tertiary influences are doing a lot most? of the work that LLM delivers. Verbatim quotes can even come from secondary sources, which themselves are using the primary material in a fair use way. In fact the high value of LLMs derives in large part because indirect sources of material give a necessary and “intelligent” context to its knowledge. In this way, the LLMs improve the primary source material, to make it truly useful. This is another way LLMs transform material. 11. When a human student is studying, they might learn from a book they purchased or from a book they borrowed from a library, which we consider fair use. We don’t judge their learning based on whether they paid for the source, or whether the sources were borrowed or rented. Likewise if an LLM is trained on borrowed copies from a library fair use that should not impact our judgement of its knowledge. 12. The constitutional rationale for US copyright is incentivizing creation for public benefit, not necessarily securing revenue streams. I’ve personally seen anecdotal evidence that LLMs have been encouraging massive degrees of co-creation. Nearly every day someone sends me something that they have recently created with the aid of an LLM. You might dismiss this as slop. Most of it is. But most of what humans traditionally made everyday could also be called slop. And the best of the slop will get better. What is clear, is that the public benefit of transforming expressions into latent space is huge, and should be encouraged. 13. At the moment the smartest AIs humans have invented are LLMs, but that might not be true in the future. There are lots of alternative models to LLMs being experimented with, and they could become the default models next. Today the most advanced LLM models are intensely commercial and rich . But they might not always be. We already have various types of open weights, open source LLMs. We tend to think of LLMs as chiefly commercial entities, trying to maximize revenue. But LLMs don’t need to be profit-making. Many of the open source models are matching the capabilities of for-profit models today. LLMs can be non-profit. They can also be considered a true common wealth, like a public domain, or like the internet. I call one version of this kind of AI, Public Intelligence. There could be more than one public intelligence. These LLMs would have multiple stakeholders, be publicly funded and publicly accountable. They could be trained on all texts in all languages from all times. In part they would be like a super library with a super librarian, able to do most knowledge work. This public intelligence would be available to anyone, at cost. The latent spaces of the LLMs whether profit or non-profit are our new public domain. The purpose of copyright is to maximize the public domain by encouraging abundance of creations with temporary monopolies and then moving the protected into the commons as soon as possible to assist in further creation. The new LLMs are a new kind of commons that provide society with new powers of understanding getting fantastic answers to questions , new ways of self-improvement, new ways to collaborate, new ways to be creative. It would be best for society and everyone born and yet unborn, if public LLM commonwealths were encouraged. We should maintain the fair use option for training LLMs to ensure we have at least some public intelligence commons.