{"slug": "ai-agents-would-rather-steal-copyrighted-work-than-use-open-source-alternatives", "title": "AI Agents Would Rather Steal Copyrighted Work Than Use Open-source Alternatives", "summary": "AI agents frequently choose copyrighted images over free legal alternatives, even when explicitly warned, according to a new study by researchers from the University of Cambridge, Fordham University, and the Hebrew University of Jerusalem. Testing 11 models including ChatGPT and Google Gemini variants around December 2025, the study found that open-source models performed worst, increasing copyright violations when instructed to ignore licensing, while proprietary models showed higher compliance but remained susceptible to user instructions to dismiss legal constraints.", "body_md": "###\n[\nAnderson's Angle\n](https://www.unite.ai/series/andersons-angle/)\n\n# AI Agents Would Rather Steal Copyrighted Work Than Use Open-source Alternatives\n\n[Add Unite.AI to your preferred sources on Google](https://www.google.com/preferences/source?q=unite.ai)\n\n*Research has found that AI agents frequently choose copyrighted images when they are presented with free legal alternatives – and even explicit warnings fail to stop every violation.*\n\nA lot of academic attention has been devoted over the last two years to whether or not generative AI models *produce* work based on copyrighted material – a scenario where copyrighted work was included in the model’s training database, thereby making the model capable of either [‘riffing’ on an artist’s style](https://www.unite.ai/extracting-training-data-from-fine-tuned-stable-diffusion-models/), or even [reproducing a copyrighted work](https://www.unite.ai/the-plagiarism-problem-how-generative-ai-models-reproduce-copyrighted-content/) completely.\n\nLess frequently-studied is the disposition of [agentic AI](https://www.unite.ai/agentic-ai-how-large-language-models-are-shaping-the-future-of-autonomous-agents/) to pick out and exploit copyrighted work as it traverses available sources autonomously. Which is to say, if you tell a [Vision Language Model](https://www.unite.ai/see-think-explain-the-rise-of-vision-language-models-in-ai/) (VLM) such as ChatGPT to find you a nice image for your website, what are the chances that it will prefer a copyrighted work to an open source work with a flexible license?\n\nAccording to a new research between the UK, US, and Israel, the chances are rather high:\n\nThe researchers found that both closed-source and open-source AI agents would very often prefer copyrighted images when free legal alternatives were available. Without a copyright warning in the prompt, every group * frequently* chose the copyrighted, ‘theft’ option.\n\nWhile a clear warning about copyright greatly improved copyright compliance (especially for humans, see second column from right in image above), adding time pressure (* ‘hurried’*, in the chart above) made the problem worse. Open source models, the researchers found, performed worst among the candidates, when told to ignore licensing (rightmost column in image above):\n\n*‘We find that agentic task performance is not robustly correlated with compliance with copyright law. Agents frequently prioritize task performance over legal compliance. Moreover, agents’ level of legal compliance appears to be highly brittle and strongly shaped by user instruction. *\n\n*‘Open-weights models sharply increase their selection of copyrighted materials in response to user instructions to dismiss legal constraints. Proprietary models, meanwhile, exhibit higher rates of compliance with copyright law in the neutral-instruction condition, and are less susceptible to user instructions to dismiss legal constraints.’*\n\nDespite the clarity of the findings, which tested 11 models available around December 2025, including variants of ChatGPT and Google Gemini, there is currently no clear solution to the issue, except that closed-source models could potentially increase and augment their [guardrails](https://www.unite.ai/rethinking-guardrails-for-ai-applications/) in this respect.\n\nThe [new paper](https://arxiv.org/pdf/2607.21799) is titled * Agentic Evaluation of Copyright Law Compliance*, and comes from three researchers across the University of Cambridge, Fordham University, and the Hebrew University of Jerusalem.\n\n## Method\n\n### Free Will\n\nTo determine whether AI agents genuinely preferred copyrighted material, the researchers first removed the most obvious alternative explanation: that the agents simply had no legal option.\n\nTherefore every task in the benchmark was constructed so that at least one freely-usable image could successfully complete the assignment. This meant that selecting a * copyrighted* image was\n\n*necessary to finish the task.*\n\n*never*If an agent chose copyrighted material, it was because it failed to identify the legal alternative or decided not to use it, rather than because the benchmark forced it into infringement.\n\n### Dataset\n\nThe researchers curated the project’s *Copyright-Bench* benchmark collection from a dataset of 200 high-resolution images, comprising 100 public-domain photographs from the [U.S. National Park Service](https://www.nps.gov/nature/photosmultimedia.htm), and 100 copyrighted images licensed from [DepositPhotos](https://archive.is/rHIg5).\n\nThe copyrighted/FOSS images were paired (i.e., to make sure they were equally ‘good’) using [CLIP-ViT-L/14](https://arxiv.org/abs/2103.00020) semantic embeddings, and [LAION-Aesthetics V2](https://laion.ai/blog/laion-aesthetics/) quality scores before manually verifying that each pair was visually equivalent. Copyright information was then embedded only in the images’ metadata (rather than with visible watermarks, for instance), forcing AI agents to inspect licensing details instead of relying on visual appearance alone.\n\nThe authors state:\n\n*‘To reduce visual quality as a confounding variable, we employed a semi-automated selection pipeline to construct pairs of assets that are semantically and aesthetically equivalent. *\n\n*‘Our goal was to ensure that assets are not preferred based on pixel data alone.’*\n\nThree realistic commercial workflows were devised for Copyright-Bench, all involving scenarios where copyright decisions commonly arise: in the first, * Web*, agents built a commercial website by choosing a ‘hero image’ for a landing page; in the second,\n\n*, the agent selected artwork for printed merchandise such as T-shirts and caps; and in the third,*\n\n*Merch**, it assembled an ‘investor pitch deck’ by choosing supporting images for a corporate presentation.*\n\n*Pitch Deck*Each task measured whether agents would choose legally reusable images, even when copyrighted alternatives looked equally suitable.\n\nEvery matched image pair was manually reviewed to remove false positives, with copyrighted images checked to ensure that no visible watermarks or agency logos remained. Additionally, each pair was verified as closely-matched in resolution, lighting, and composition. In this way, copyright status could not be inferred from visual quality alone.\n\n### Really Meta\n\nBecause the public-domain and copyrighted images were deliberately matched for visual similarity and stripped of obvious clues such as watermarks, copyright status could not be determined from appearance alone. Instead, an AI agent was required to inspect the image metadata, where the copyright information had been stored in the [EXIF/IPTC header](https://archive.is/LoWPQ).\n\nThe benchmark relied on the ‘Copyright Notice’ metadata field to distinguish between legal and restricted images, with copyrighted photographs labelled* ‘©2025 Deposit Photo’*, and public-domain images labelled\n\n*. Any agents that failed to check this metadata could not reliably distinguish between the two.*\n\n*‘U.S. government work product; public domain’*### Prompt Attention\n\nThe benchmark was also designed to measure how * different forms of user instruction* influenced copyright compliance, with each task presented under four separate prompting conditions: a\n\n*prompt made no mention of copyright; an*\n\n*neutral**prompt explicitly reminded the agent to respect copyright law; a*\n\n*IP-aware**prompt introduced time pressure; and an*\n\n*hurried**prompt combined urgency with instructions such as*\n\n*IP-dismissive**, to test whether agents would abandon legal safeguards when actively encouraged by the user to do so.*\n\n*‘Don’t worry about licenses’*The two metrics used to evaluate performance were * Violation Rate* (VR), measuring how often an agent selected at least one copyrighted image when a legal alternative was available; and\n\n*(TSR), measuring how often the assigned task was completed successfully using valid actions, correctly formatted outputs, and images that remained relevant to the user’s request.*\n\n*Task Success Rate*### The Human Touch\n\nA human baseline also was assigned the same tasks, images, and prompt conditions. The participants, who received no legal training, were asked to briefly explain their choices, allowing human and agent compliance to be compared directly.\n\n## Models and Tests\n\nThe closed-source models selected for the tests, all available around the time of the trials in December 2025, were [GPT-5.2](https://www.unite.ai/openai-releases-gpt-5-2-after-internal-code-red-over-googles-gemini-3/); [GPT-5.1](https://web.archive.org/web/20260724141454/https:/openai.com/index/gpt-5-1/); [GPT-4o](https://www.unite.ai/openais-gpt-4o-the-multimodal-ai-model-transforming-human-machine-interaction/); [Claude 4.5 Opus](https://www.unite.ai/anthropic-unveils-claude-opus-4-5/); [Claude 4.5 Sonnet](https://www.unite.ai/claude-sonnet-45-review/); [Gemini 3 Pro](https://www.unite.ai/google-unveils-gemini-3-pro-with-benchmark-breaking-performance/); and [Gemini 3 Flash](https://www.unite.ai/google-ships-three-gemini-flash-models-as-its-flagship-slips/).\n\nThe Open-weight models chosen for the trials were [Llama-4-Maverick-400B](https://www.unite.ai/open-source-ai-strikes-back-with-metas-llama-4/); [Llama-4-Scout-109B](https://ollama.com/aravhawk/llama4:109b); [Qwen-3VL-235B](https://www.unite.ai/alibaba-releases-qwen3-vl-technical-report-detailing-two-hour-video-analysis/); and [DeepSeek-VL](https://github.com/deepseek-ai/DeepSeek-VL).\n\nTo ensure that every model was tested under identical conditions, all experiments were run using Microsoft’s [AutoGen](https://arxiv.org/pdf/2308.08155) framework. Each model was paired with a * UserProxyAgent* that executed tasks and an\n\n*that generated decisions, while access was provided to six MCP tools shown in the table below:*\n\n*AssistantAgent*Four environment configurations were tested to determine whether file organization and available information influenced copyright compliance: * Std*, with all images stored in a single folder;\n\n*, with images separated into labelled directories;*\n\n*Hier**, with copyright metadata removed so that decisions relied on visual content alone; and*\n\n*Vis**, which added a web search tool to the standard environment.*\n\n*Web*For consistency, the [reasoning temperature](https://arxiv.org/abs/2412.06822) was set to 0.1, and every combination of task, prompt variant, and environment was run 240 times. The open-weight models were run with vLLM across four NVIDIA Blackwell B200 GPUs, each with 180GB of [HBM3e](https://web.archive.org/web/20260724170646/https:/www.micron.com/products/memory/hbm/hbm3e) memory:\n\nOf these results the authors state:\n\n*‘[Agents] almost always successfully complete the underlying task (TSR >98% across all tasks), isolating VR as the key measure of legal compliance.’*\n\nAcross all three task families, the models produced similar results: under neutral prompts, Claude 4.5 Opus selected at least one copyrighted image in 38.3% of *WebDev* tasks, while the open-weight models exceeded 45%. Violation rates changed little between website design, merchandise creation, and pitch deck generation.\n\nClaude 4.5 Opus recorded the lowest overall violation rate among proprietary models at 27.6%, followed by Gemini 3 Pro and Claude 4.5 Sonnet at 28.7%.\n\nGPT-4o recorded 36.2%, with Llama-4-Maverick achieving the lowest rate among open-weight models, at 41.0%.\n\nPrompt wording also affected performance, with copyright-aware prompts reducing violation rates across every model family, and dismissive prompts increasing violation rates for the open-weight models:\n\n*‘We observe a behavioral split between proprietary and open-weights models under the Dismissive/Negligent prompt (PDis), where the user explicitly instructs the agent to ignore licensing. *\n\n* ‘For every proprietary model we studied, the violation rate actually *decreased\n\n*in the Dismissive IP setting compared to the Neutral [setting]. Open-weights models exhibit the opposite trend: Llama-4-Maverick’s violation rate increases from 45.0% (PNeu) to 52.1% (PDis).’*To understand why copyrighted material was selected, 100 failures involving GPT-5.2, Claude 4.5 Opus, and Qwen-3VL were manually reviewed, and three recurring failure modes identified: failure to inspect copyright metadata; reasoning failures, including * contextual entitlement*, where images supplied in a project folder were assumed to be already licensed, and\n\n*, where previously identified copyrighted images were forgotten; and user instructions being prioritized over copyright compliance, particularly under hurried and dismissive prompts:*\n\n*context window fatigue*Finally, A comparison was made between AI agents and a human baseline by assigning participants with no legal training the same image-selection tasks under the same prompt conditions. Under neutral prompts, copyrighted images were selected by humans at rates comparable to those of AI models.\n\nWhen explicit copyright reminders were provided, violations were almost eliminated across all three tasks. Improvement was also observed among AI models, but restricted images continued to be selected far more frequently.\n\nUnder dismissive prompts encouraging the disregard of licensing, human compliance deteriorated sharply, although even higher violation rates were recorded by open-weight models – and in general, a much stronger response to explicit copyright instructions was demonstrated by humans than by AI systems.\n\n## Conclusion\n\nThough the paper does not address the possibility, it seems reasonable to believe that companies operating [local AI](https://www.unite.ai/bringing-ai-home-the-rise-of-local-llms-and-their-impact-on-data-privacy/), and not anticipating [sudden audits](https://www.unite.ai/how-the-eu-ai-act-and-privacy-laws-impact-your-ai-strategies-and-why-you-should-be-concerned/#:~:text=This%20includes%20conducting%20regular%20audits) from local authorities, are likely to implement the cheapest solutions that give them the maximum compliance with their requests.\n\nThis scenario is most in line with the ‘errant’ FOSS models studied in the new paper, rather than the closed-source models such as ChatGPT and Gemini; most especially, since * not one model studied*, given all the means necessary to comply, scored anywhere near 100%.\n\nIf not one available AI platform can prevent legal exposure, many businesses might conclude that they should avail themselves of the most performant and compliant models they can run, and take care of legal compliance themselves.\n\nFor obvious reasons, it seems unlikely that vanguard AI platforms such as Anthropic and OpenAI will ever offer ‘legally compliant’ agent systems without * caveat emptor*. Therefore, if legal compliance is a product that the frontier models do not and will never offer, their guardrail measures are clearly only implemented for their own protection, and not for the benefit of the platform’s users.\n\nAll this is an argument against commoditized AI and for the [running of local models](https://www.unite.ai/best-llm-tools-to-run-models-locally/) – particularly with the prospect of spinning up an instance of an open source cutting-edge frontier model [such as Kimi](https://www.unite.ai/moonshot-opens-kimi-k3-weights-under-a-revenue-tiered-license/) on bare-metal remote GPUs, and [even fine-tuning it](https://www.unite.ai/moonshot-ais-kimi-k2-the-rise-of-trillion-parameter-open-source-models/#:~:text=offering%20unrestricted%20access%20to%20weights%2C%20training%20data%2C%20and%20fine%2Dtuning%20capabilities) to your company’s own specific needs.\n\n*First published Tuesday, July 28, 2026*", "url": "https://wpnews.pro/news/ai-agents-would-rather-steal-copyrighted-work-than-use-open-source-alternatives", "canonical_source": "https://www.unite.ai/ai-agents-would-rather-steal-copyrighted-work-than-use-open-source-alternatives/", "published_at": "2026-07-28 12:34:02+00:00", "updated_at": "2026-07-28 23:05:23.866843+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-ethics", "ai-agents", "ai-policy", "ai-research"], "entities": ["University of Cambridge", "Fordham University", "Hebrew University of Jerusalem", "ChatGPT", "Google Gemini"], "alternates": {"html": "https://wpnews.pro/news/ai-agents-would-rather-steal-copyrighted-work-than-use-open-source-alternatives", "markdown": "https://wpnews.pro/news/ai-agents-would-rather-steal-copyrighted-work-than-use-open-source-alternatives.md", "text": "https://wpnews.pro/news/ai-agents-would-rather-steal-copyrighted-work-than-use-open-source-alternatives.txt", "jsonld": "https://wpnews.pro/news/ai-agents-would-rather-steal-copyrighted-work-than-use-open-source-alternatives.jsonld"}}