cd /news/ai-policy/the-government-is-buying-ai-by-the-t… · home › topics › ai-policy › article
[ARTICLE · art-146183] src=fedscoop.com ↗ pub= topic=ai-policy verified=true sentiment=↓ negative

The government is buying AI by the token. It should buy results.

Federal agencies began buying ChatGPT through a consumption-based token model this month under the OneGov agreement between GSA and OpenAI, which replaced the prior $1-per-agency-per-month charge with a 50% discount and a 27-month term covering approximately 23 million public servants, according to GSA acting Federal Acquisition Service commissioner Laura Stanton. The token model conflicts with executive order 14402, issued in April, which made fixed-price contracts with performance-based considerations the "default and preferred method of procurement," and with the Pentagon CDAO's June selection of Accenture Federal Services for an $821 million War Data Platform integration task order designed to deliver quantifiable outcomes. The author, a former Department of Defense customer experience officer, argues token pricing rewards volume rather than results and prevents the government from measuring cost-effectiveness.

by read7 min views1 publishedOct 6, 2026
The government is buying AI by the token. It should buy results.
Image: Fedscoop (auto-discovered)

Starting this month, federal agencies began purchasing ChatGPT through the token. The recently signed OneGov agreement between GSA and OpenAI replaced the previous model of charging each agency $1 per month with a 50% discount for a consumption-based payment model.

It will last for 27 months, and according to OpenAI, create access to the program for approximately 23 million public servants across all levels of government. “This is the next logical step to providing access to the platform,” stated Laura Stanton, GSA’s acting commissioner of Federal Acquisition Service.

While the consumption-based model is logical from a vendor’s perspective and can help increase revenues, it does not allow the government to measure its cost-effectiveness and goes against the government acquisition policy. When evaluating a service during my five years working in the Department of Defense — first in the Defense Digital Service and later as the DOD’s customer experience officer — I had to assess how much users had benefited from it. A usage metric would not be able to answer my question.

The rest of the government points the other way #

Issued in April, executive order 14402 said “fixed-price contracts with performance-based considerations” would be deemed the “default and preferred method of procurement.”

The order characterized fixed-price contracts as having “clearly defined outcomes and deliverables on predictable timelines” that “often tie profit to the contractors’ performance.” Written justification is now required for any cost-reimbursement contract above certain thresholds.

That approach is already being applied to data and AI by the Pentagon’s Chief Digital and AI Office. In June, the GSA selected Accenture Federal Services for an $821 million integration task order for the War Data Platform through the CDAO. A Pentagon official told DefenseScoop that the task order was “explicitly designed to deliver measurable, high-impact mission outcomes that can be quantified (not time-based deliverables).”

According to that official, the office planned to “aggressively pursue the transition of the contract into a fixed-price vehicle once we have completed the first year of the hybrid effort.”

A former defense official, however, told DefenseScoop that the approach appeared inconsistent with the solicitation itself. I would have initiated it in a fixed-price arrangement.

That’s the right argument to have, and it’s missing from the market for the models themselves. The Pentagon’s frontier AI awards from last year, capped at $200 million each, were prototype agreements. GSA’s governmentwide deal now charges by the token. In September, CDAO started a shared-savings pilot with Red Cell Partners that pays AI vendors only for the costs they eliminate. The pilot is a real step for AI services, but I haven’t found an agreement that prices model access on the result it produces.

Token pricing rewards volume #

Appearing on CNBC in July, Palantir CEO Alex Karp said “something has gone completely wrong” with how the labs sell tokens. He described the basic view among U.S. enterprises as wasting time on tokens, getting no value and handing the labs their intellectual property.

Karp is no more neutral than I am, since Palantir sells platforms and a platform vendor gains when model access becomes a commodity. The incentive he describes is still real.

A per-token contract pays the vendor more when a workflow uses more tokens. Longer prompts raise the bill, and so do longer answers and agents that loop through extra steps. None of it tells the Pentagon whether an aircraft that should have flown did fly.

The contract rewards use and treats the model like office software. The value comes from building failure prediction into how Air Force maintainers order parts and schedule repairs. A chatbot on every maintainer’s laptop doesn’t change how either one works.

Nobody measures what the users get #

Outcome pricing needs a measured outcome, and here the federal market has a gap. The public numbers on government AI are usage numbers: GenAI.mil has about 1.7 million users, some 500,000 of them daily “power users.” The workforce has built more than 100,000 agents on it. Those are strong adoption numbers, and the team behind GenAI.mil deserves credit for them. They say nothing about whether the answers were right.

Much of that use also runs through integrators and platforms that sit between the lab and the agency. The lab sees tokens and the integrator sees its own application. The agency gets an invoice, and no party in that chain is responsible for checking the answer.

CNN reported last month that a chatbot misidentified a Chinese vessel’s cargo as components for a nuclear weapons program. The special operations command analyst who queried it then used AI again to turn the finding into a formal intelligence report. Armed personnel prepared to board the ship, and aircraft were airborne before officials looked closely at the report. A per-token contract bills the tokens behind that report at the same rate as the tokens behind a correct one.

Outcomes are harder to buy than tokens #

Replacing token pricing with outcome pricing isn’t a clean fix.

First, someone has to define the outcome. “Failures predicted” can mean more flags, which pulls good parts off aircraft and adds work for maintainers. A vendor paid per flag has the same bad incentive as a vendor paid per token, but with a different unit.

Second, the vendor can’t grade its own work. If the lab or the integrator reports the outcome metric, the contract ends up measuring their reporting.

Third, outcomes take time to observe. A maintenance model’s value shows up in mission-capable rates months later, but the token invoice arrives every month. Program offices under budget pressure will pick the metric they can see.

Staying on token pricing doesn’t solve any of those problems, so I’d build the measurement first and follow these steps:

Put evaluation in the contract. Every AI task order above a threshold should name the mission metric, the baseline and who measures it. The measuring party should be independent of the vendor and the integrator. EO 14402 already asks for “clearly defined outcomes,” and for AI, that definition is an evaluation plan.

Use the workforce as the evaluator. Federal employees already know which answers are wrong. An analyst who catches a bad cargo identification produces the most valuable data in the system, and today that judgment goes nowhere. Agencies should capture expert judgment on model output in a structured way and at scale, and the contract should say the agency owns it.

Fund testing outside the labs. Scale AI CEO Francis deSouza argued in a blog post last month that most frontier model testing happens inside the labs and that Washington needs its own evidence on cyber, biological and other national security risks. Scale works with government AI evaluation bodies in four countries and gains from more testing, but the funding figures support his point. For fiscal 2027, the president requested $27 million for the Center for AI Standards and Innovation, the government’s frontier AI testing center, and House appropriators proposed up to $15 million.

Pilot fixed-price AI where results are countable. I’d start with aircraft mission-capable rates, depot repair times and spare-parts backorders, which already have baselines.

CDAO and the undersecretary for research and engineering, who oversees it, have shown they will push a large data contract toward fixed price. They should do the same for model access, with firm fixed-price or shared-savings terms tied to a measured result. Intelligence analysis can come after that works.

GSA’s agreement runs for 27 months. That’s enough time to run a fixed-price maintenance pilot and measure what it does to mission-capable rates. Whoever negotiates the next AI agreement should have that number in hand.

That number will exist only if someone measures it, and measuring costs little next to what is being bought. The $27 million the president requested for CAISI is less than a seventh of the $200 million ceiling on any one of the Pentagon’s frontier AI awards.

Agencies that want to know before the 27 months are up whether the answers were right need to write the evaluation into their task orders now.

Savan Kong is a co-founder of Your Roster. He previously worked at the Defense Digital Service and served as the Defense Department’s customer experience officer.

── more in #ai-policy 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/the-government-is-bu…] indexed:0 read:7min 2026-10-06 · —