OpenAI’s Project Giraffe: Tracking Copyright in 78M Chats, Then Deleting the Logs OpenAI secretly built 'Project Giraffe,' an internal evaluation suite that used probabilistic Bloom filters to monitor copyright infringement across 78 million de-identified user conversations, according to leaked deposition testimony in the OpenAI copyright lawsuit. The company then deleted the output logs, transforming routine database hygiene into potential federal evidence spoliation. The revelation dismantles OpenAI's previous defense that searching ChatGPT training corpora for copyrighted content was technically impossible. Member-only story OpenAI’s Project Giraffe: Tracking Copyright in 78M Chats, Then Deleting the Logs Leaked deposition testimony in the landmark OpenAI copyright lawsuit reveals how engineers deployed probabilistic Bloom filters to monitor plagiarism across 78 million user conversations: and why the subsequent deletion of those files transforms routine database hygiene into federal evidence spoliation. For more than two years, OpenAI told federal magistrates that searching ChatGPT training corpora for copyrighted news or auditing user completions was technically impossible. Leaked deposition testimony from an April discovery hearing in the OpenAI copyright lawsuit just revealed the company secretly built “Project Giraffe”: an internal evaluation suite that used probabilistic Bloom filters across 78 million de-identified user conversations to measure verbatim copyright infringement before allegedly deleting the output logs. The algorithmic mechanics of Project Giraffe dismantle the “unsearchable black box” defense completely: leaving enterprise AI engineers with an urgent mandate to restructure their data preservation protocols before internal safety telemetry becomes Exhibit A in intellectual property…