cd /news/artificial-intelligence/five-ways-to-run-the-same-ai-job-thr… · home › topics › artificial-intelligence › article
[ARTICLE · art-148418] src=dev.to ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Five Ways to Run the Same AI Job, Three Kinds of Mistakes, and What Broke at 100,000 Records

A developer benchmarked five ways of running the same AI classification job on fake customer support chats — four batch API services (OpenAI, Gemini, Claude with thinking on and off) plus a self-hosted model — and found that while all five nailed product and problem-type labels, they diverged sharply on sentiment, with OpenAI and the self-hosted model skewing too negative and Gemini too forgiving. Scaling the job to 100,000 chats across the three batch services produced no missing answers or retries, but latency varied wildly: Gemini finished in about 45 minutes, Claude took 2 hours 44 minutes, and OpenAI's per-chunk time stretched from 14 minutes to nearly 90 minutes depending on time of day, at roughly 9 cents per 1,000 chats for OpenAI and Gemini and 5 cents for Claude. The self-hosted model processed 1,000 chats in 29 seconds for about $0.004 per 1,000 chats, though it was the least accurate on sentiment and required roughly 8 minutes of setup on a rented GPU.

by read5 min views3 publishedOct 9, 2026

Originally published at kotwal-itpro.github.io. Two days ago I wrote about giving two AI models the same instructions and watching them get things wrong in opposite ways. Since then I've added two more ways of running the job, made it 100 times bigger, and run it again on different days. Here's what changed, and what didn't.

Same job as before. I have fake customer support chats, and for each one the AI has to say which product it's about, what kind of problem it is, how the customer feels, and whether they asked to be contacted again. I wrote the chats, so I know every right answer.

This time I ran it five ways:

The first four are "batch" services. You send everything at once and collect the answers later, for about half the usual price. The last one has no queue at all. It's just a computer doing the work.

All five got the product and the problem type right every time. Nearly all of them got the follow-up question right too, once I'd explained it properly (that was the big lesson last time).

The "how does the customer feel" question is where they split, into three groups:

Too negative. OpenAI (93% right) and my self-hosted model (85% right). A customer writes "My laptop stopped working after three days. Let me know what you need from me." That's a calm, helpful message about a problem. Both called the customer unhappy. My own model did this twice as often, and sometimes even turned happy customers into unhappy ones.

Too forgiving. Gemini (82% right). "This is really frustrating" came back as neutral.

Depends on one setting. Claude scored about the same both ways, 89% and 91%. With thinking on, its mistakes were split about evenly between too harsh and too forgiving. With thinking off, almost all of them were too harsh, like OpenAI's. Same model, same instructions, one switch, and the balance of its mistakes changed. The two settings gave different answers on about 1 chat in 9.

If you take one thing from this post: a score doesn't tell you which way a model is wrong. If something downstream counts unhappy customers, two models with similar scores can give you very different counts. The self-hosted model surprised me. Once it was set up, it did all 1,000 chats in 29 seconds. The hosted services took 3 to 10 minutes, because you wait in their queue. It also cost about $0.004 per 1,000 chats in GPU time, about a tenth of the cheapest hosted option.

The catch is the setup. Installing the software, down the model and getting it ready took about 8 minutes on the rented GPU, which cost more than the actual work. And it was the least accurate of the five on feelings. It's also an older, smaller model than the hosted ones, so that isn't the last word on self-hosting.

Then I made the job 100 times bigger: 100,000 chats, on each of the three batch services.

Nothing went missing. All three sent back all 100,000 answers. None needed a retry, and the scores matched the smaller runs almost exactly. I'd half expected the trouble to start here. It didn't.

The waiting changed a lot. Gemini did it in 10 pieces of 10,000, about 4 minutes each, so 45 minutes in total. Claude took all 100,000 in one go and needed 2 hours and 44 minutes. OpenAI was the odd one. Its pieces took 14 minutes early in the afternoon, then slowed to almost an hour and a half by the evening. Same job, same size, six times slower depending on when I sent it.

The price per 1,000 chats didn't change with size: about 9 cents on OpenAI and Gemini, 5 cents on Claude.

The only real failure was on my side. Eight hours into the OpenAI run, my internet connection dropped while up the ninth piece, and the job stopped. This is exactly what I built the tool for. It keeps a record of what's done, so when I ran the same command again it picked up the last 20,000 chats and didn't resend, or pay for, the 80,000 that were already finished. It now also waits and retries when the connection drops, instead of giving up.

One thing went wrong before any work started. I sent Gemini all 100,000 chats as one job. It took five minutes to upload, then Gemini said:

You exceeded your current quota, please check your plan and billing details.

That sounds like a billing problem. It wasn't. When I sent the same work in pieces of 10,000, one after another, every piece went straight through. The real problem was that one job was too big for how much work my account can have waiting. The message just didn't say so.

That one changed my tool. It now recognizes a "queue full" answer, waits or sends smaller pieces, and doesn't count it as a failed attempt for the records involved.

This was the question I most wanted answered. I sent the exact same 1,000 chats, with the same instructions, to each service on three days in a row: Wednesday, Thursday and Friday.

Here's the strange part. The overall scores barely moved from day to day: Gemini's "how does the customer feel" score was 82%, 81% and 81%. But underneath, a different set of chats was right each day. Every single chat that changed was answered correctly on one day and wrongly on another. The services wobble on the borderline cases, and which side a case lands on changes from day to day.

If you compare this week's numbers with last week's, some of the change is the model, not your customers. Know how big that wobble is before you read anything into a small shift. I've written these up properly in a guide. The tool is now on PyPI (pip install steadybatch), and the code, the fake chats and every result are on GitHub at kotwal-itpro/steadybatch. If you want to cite it, it now has a DOI: 10.5281/zenodo.23221956.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/five-ways-to-run-the…] indexed:0 read:5min 2026-10-09 · —