A 6-Case Single API Key Acceptance Harness for Compatible SaaS Chat A developer describes a six-case acceptance harness for scoring job candidates through a single API key across OpenAI, Claude, and Gemini, arguing that transport compatibility does not guarantee behavioral equivalence. The approach defines a versioned scoring contract around a scoreCandidate() boundary, validates every provider response locally, preserves null for unknown evidence, and computes weighted scores in application code. Failures that exhaust the retry budget are routed to manual review rather than silently converted to a zero. A single API key saves deployment work, but it does not make model behavior portable. For a SaaS that scores candidates against a job rubric, the useful unit of portability is a versioned scoring contract backed by six acceptance cases. My choice is one small adapter boundary, validation on every response, and promotion only after a provider-model pair clears that harness. Short answer: compare integrations by the application behavior they preserve, not the credentials they replace. OpenAI, Claude, and Gemini can sit behind one application interface, yet each remains a distinct execution target. A compatible chat-shaped request is transport compatibility. Stable candidate decisions are a product requirement. This matters for a one-person SaaS. Saving an hour during setup helps once. Avoiding silent scoring changes protects every weekly release. I outsource secret storage and schema validation because those are undifferentiated; the rubric and decision rule stay in application code. One credential answers how a service authenticates. Candidate scoring raises harder questions. Did every rubric item appear? Is each score tied to evidence? Does absent evidence become unknown , or does the model invent certainty? Those questions cannot be settled when an endpoint accepts a request. Supporting OpenAI, Claude, and Gemini through a common request shape does not imply identical outputs. I would record each exact model as a tested target and never infer behavioral equivalence from a provider name. The constraint that changes the design is simple: untrusted resume text enters upstream, while a hiring workflow consumes the result downstream. A malformed summary is annoying. A plausible, unsupported score can affect who receives review. The first abstraction should therefore be scoreCandidate , not a generic chat method. Keep that boundary dull. Batch work also deserves a separate lifecycle. The OpenAI Batch API guide describes uploaded input, batch creation, status checks, and output retrieval. That is not an interactive call. A portable application should model deferred jobs separately instead of forcing both paths through one blocking chat abstraction. The contract carries domain facts rather than provider vocabulary. It preserves null for unknown evidence. Zero means a criterion failed; null means the submitted material cannot support a decision. Combining them creates false precision. type Criterion = { id: string; description: string; weight: number }; type ScoreRequest = { rubricVersion: string; criteria: Criterion ; candidateText: string; }; type CriterionResult = { criterionId: string; score: 0 | 1 | 2 | null; evidence: string ; }; type ScoreResult = { rubricVersion: string; results: CriterionResult ; }; interface ScoringTarget { id: string; score input: ScoreRequest : Promise