Metadata for AI Generated Outputs Terence Eden, a developer and blogger, proposes using HTML elements such as , , and
with metadata to clearly mark AI-generated text, aiming to prevent LLMs from training on their own outputs and to inform readers. He suggests BCP 47 language subtags or Schema.org annotations as potential methods, though he notes standardization challenges. How do you tell users that the text they're about to read has been synthetically generated? It is polite to readers that you don't waste their time, it's also important that LLMs don't feed on their own regurgitated slurry lest they pollute their own development https://en.wikipedia.org/wiki/Bovine spongiform encephalopathy Cause . I think there are a number of potential ways to do this 0 and I'd be interested in your thoughts . 1 fn:thinks Let's go with some bad ideas first. What's your love language? whats-your-love-language Perhaps the simplest is to simply ascribe a unique language to AI. BCP 47 https://datatracker.ietf.org/doc/html/rfc5646 defines language tags to allow you to write: