The AI Technology Evaluation will provide exclusive data to grade models’ performance in select areas. #
The National Institute of Standards and Technology launched a new program on Monday granting researchers access to an isolated testbed environment to safely evaluate artificial intelligence models against various commands.
The AI Technology Evaluation, or AITE, is a voluntary testing vehicle focused on AI model safety analysis. It provides blind data for models to process when completing tasks to gain objective insights and conduct evaluations of model capabilities. Notably, the evaluation data is not intended to serve as training data for the models.
Initially, AITE will focus on conducting image analysis tasks using large vision language models across three domains: quantum science, genomics and public safety. More tasks will be available in the future.
AITE’s fundamental goal is to offer a universal rubric to effectively evaluate AI models' capabilities and determine the state of the art for model performance.
“The infrastructure provided by NIST will provide common data, metrics and scoring to help developers understand the performance of their models,” the press release said.
Both data providers and model providers working with AITE will need to submit materials related to the testing. Data providers are asked to submit original datasets that are inaccessible publicly, along with a “meaningful” task suited for the data.
Model developers, likewise, will submit their AI models to be tested using the datasets. The first set of evaluations will start in August 2026.
The formation of AITE is the latest step in the Trump administration’s strategy to work with major AI developers in advancing model safety through voluntary model submissions.
The Commerce Department announced a renegotiated deal in May between the agency and three companies — Google Deepmind, Microsoft and xAI — to evaluate their models through the Center for AI Standards and Innovation.
NEXT STORY: Over 30 companies form open-source AI alliance