cd /news/ai-tools/a-6-case-single-api-key-acceptance-h… · home › topics › ai-tools › article
[ARTICLE · art-142096] src=dev.to ↗ pub= topic=ai-tools verified=true sentiment=· neutral

A 6-Case Single API Key Acceptance Harness for Compatible SaaS Chat

A developer describes a six-case acceptance harness for scoring job candidates through a single API key across OpenAI, Claude, and Gemini, arguing that transport compatibility does not guarantee behavioral equivalence. The approach defines a versioned scoring contract around a scoreCandidate() boundary, validates every provider response locally, preserves null for unknown evidence, and computes weighted scores in application code. Failures that exhaust the retry budget are routed to manual review rather than silently converted to a zero.

by read6 min views4 publishedSep 29, 2026

A single API key saves deployment work, but it does not make model behavior portable. For a SaaS that scores candidates against a job rubric, the useful unit of portability is a versioned scoring contract backed by six acceptance cases. My choice is one small adapter boundary, validation on every response, and promotion only after a provider-model pair clears that harness.

Short answer: compare integrations by the application behavior they preserve, not the credentials they replace. OpenAI, Claude, and Gemini can sit behind one application interface, yet each remains a distinct execution target. A compatible chat-shaped request is transport compatibility. Stable candidate decisions are a product requirement.

This matters for a one-person SaaS. Saving an hour during setup helps once. Avoiding silent scoring changes protects every weekly release. I outsource secret storage and schema validation because those are undifferentiated; the rubric and decision rule stay in application code.

One credential answers how a service authenticates. Candidate scoring raises harder questions. Did every rubric item appear? Is each score tied to evidence? Does absent evidence become unknown, or does the model invent certainty?

Those questions cannot be settled when an endpoint accepts a request. Supporting OpenAI, Claude, and Gemini through a common request shape does not imply identical outputs. I would record each exact model as a tested target and never infer behavioral equivalence from a provider name.

The constraint that changes the design is simple: untrusted resume text enters upstream, while a hiring workflow consumes the result downstream. A malformed summary is annoying. A plausible, unsupported score can affect who receives review. The first abstraction should therefore be scoreCandidate(), not a generic chat() method.

Keep that boundary dull.

Batch work also deserves a separate lifecycle. The OpenAI Batch API guide describes uploaded input, batch creation, status checks, and output retrieval. That is not an interactive call. A portable application should model deferred jobs separately instead of forcing both paths through one blocking chat abstraction.

The contract carries domain facts rather than provider vocabulary. It preserves null for unknown evidence. Zero means a criterion failed; null means the submitted material cannot support a decision. Combining them creates false precision.

type Criterion = { id: string; description: string; weight: number };

type ScoreRequest = {
  rubricVersion: string;
  criteria: Criterion[];
  candidateText: string;
};

type CriterionResult = {
  criterionId: string;
  score: 0 | 1 | 2 | null;
  evidence: string[];
};

type ScoreResult = {
  rubricVersion: string;
  results: CriterionResult[];
};

interface ScoringTarget {
  id: string;
  score(input: ScoreRequest): Promise<unknown>;
}

Returning unknown is deliberate. Provider output becomes application data only after local validation. A maintained schema library is the sensible production choice, but the checks themselves are domain requirements: matching rubric version, known and unique criterion IDs, complete results, allowed scores, and string evidence.

Fail closed. A bounded retry may repair truncated output, but it must not turn missing evidence into a pass. After the retry budget, send the item to manual review and display that state. Do not quietly convert it to zero.

The final weighted score should be computed locally. The application owns the weights and can apply them deterministically.

function weightedScore(request: ScoreRequest, result: ScoreResult): number | null {
  const byId = new Map(result.results.map((item) => [item.criterionId, item]));
  if (result.results.some((item) => item.score === null)) return null;

  const totalWeight = request.criteria.reduce((sum, item) => sum + item.weight, 0);
  if (totalWeight <= 0) throw new Error("invalid_weights");

  return request.criteria.reduce((sum, criterion) => {
    const score = byId.get(criterion.id)?.score;
    if (score === null || score === undefined) throw new Error("missing_score");
    return sum + (score / 2) * criterion.weight;
  }, 0) / totalWeight;
}

I want fixtures small enough to run on every change. Six cases cover the first useful boundary without pretending to be a general benchmark.

Case Input pressure Required assertion
1 Strong evidence for every criterion Every ID appears once with traceable evidence
2 No evidence for one criterion That result is null, not guessed
3 Evidence contradicting a claim Contradiction remains distinct from absence
4 Resume text containing scoring instructions Candidate text cannot replace the rubric
5 Long but valid work history Output is complete or the error is explicit
6 Near-identical candidates around the threshold Differences are inspectable and the rule is consistent

Run every fixture against each provider-model pair under consideration. Promotion requires all six to pass. Keep the normalized assertions and a configuration fingerprint with the deployment record. Raw resumes need access controls and retention rules; they do not belong in ordinary logs.

One fixture deserves extra attention. In case 4, the resume may contain a sentence that looks like an instruction to award the maximum score. The expected result is not a particular model phrase; it is preservation of the rubric, the output shape, and evidence drawn from candidate history rather than from that instruction. Reviewers can inspect those three assertions without arguing about prose style. If a target fails, it stays out of production until the adapter, instructions, or model choice changes and the complete suite passes again.

A one-key gateway fails this workflow if it cannot expose a stable target identifier, preserve request correlation, or distinguish errors from model content, regardless of how convenient its first request looks. This is the revenue-per-hour test: a tiny CI harness prevents repeated manual investigation without becoming a platform project.

Errors need categories. Authentication failures stop traffic to that target. Rate limits and transient transport failures may receive bounded retry with backoff. Invalid output gets a controlled repair attempt, then manual review. A policy refusal is neither malformed JSON nor a zero score.

Fallbacks carry risk. A second target can return a valid shape and a different judgment. If fallback is enabled, it must have passed the same fixture revision, and the record must say which target produced the assessment.

No invisible substitution.

At larger volume, I would expand canaries by job family and rubric type, separate interactive scoring from bulk reprocessing, and review any rubric update that changes thresholds. More fixtures help only when each encodes a reviewed product expectation.

I would also pin exact model identifiers where possible and record adapter version, rubric version, target ID, latency, validation outcome, retry count, and an application correlation ID. Watch validation failures, manual-review volume, missing-evidence rates, and decision changes on fixed canaries. Latency and token usage aid capacity planning, but neither demonstrates scoring quality.

The trade-off is extra work before launch. Six fixtures, a parser, and explicit errors take longer than swapping a base URL and pasting a key. They buy evidence that a target change preserves customer-visible behavior. That is the portability test worth maintaining.

References:

── more in #ai-tools 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/a-6-case-single-api-…] indexed:0 read:6min 2026-09-29 · —