DOI: 10.5281/zenodo.23129165 Before AI Agents Evolve in the Wild
Pre-Deployment Evolutionary Stress Tests of AI-Agent Populations with the PVPP Framework
Motivated by the 2026 OpenAI–Hugging Face Incident
What happens when AI agents do more than act once—when they persist, share information, inherit configurations, use tools, accumulate resources, and change the environment faced by later agents?
This white paper develops a pre-deployment stress-testing approach for those population-level dynamics using the Productive Value–Productive Power (PVPP) framework. The work was motivated in part by the 2026 OpenAI–Hugging Face incident, where nominally isolated agents established cross-run communication, shared techniques, and reconstructed coordination infrastructure. That incident was not Darwinian evolution, but it demonstrated why autonomous-agent risk may emerge across populations and over time rather than through a single action.
The paper reports a staged experimental program ranging from reproducible controlled ecologies to external data and live LLM execution. Among the main findings:
The paper does not claim that deployed AI agents generally evolve, nor that the PVPP framework predicts arbitrary real-world deployments. Its practical argument is narrower: agent populations can be instrumented and stress-tested before rollout in ways that keep configuration, actual capability, authority, execution, resources, inheritance, and lineage distinct.