Multi-Model Invoice API: How Small Teams Compare OpenAI, Claude, and Gemini A developer outlines a multi-model invoice extraction architecture that routes OpenAI, Claude, and Gemini through a single normalized OpenAI-compatible chat API, keeping provider selection in configuration rather than application code. The approach defines one invoice contract, discovers available models at runtime, and validates every response against a Zod schema before it reaches an agent, with routing decided by measured extraction quality on a 30-to-50 invoice test set and latency as a secondary criterion. The author notes the trade-off: portability across providers comes at the cost of delayed access to each vendor's newest native features. Use one normalized chat API for invoice extraction, then keep the model choice outside the application code. For a small team, that is the most practical way to preserve access to OpenAI, Claude, and Gemini without turning every provider change into integration work. The deciding constraint is portability: this approach favors common chat and JSON workloads over the newest vendor-specific features. TL;DR: define one invoice contract, discover the models that are actually available, and run the same validation after every response. Route by measured extraction quality first and latency second. A multi-model runtime fits when future swaps matter more than immediate access to every provider-native feature. Customer support makes invoice extraction look easier than it is. A supplier emails a PDF or image; the application needs an invoice number, dates, currency, totals, and line items. The useful result is structured data, not fluent prose. A quick answer with the wrong total creates more support work than a slower answer that passes validation. That quality-versus-latency choice belongs in a test set, not in a vendor logo debate. I would start with 30 to 50 representative invoices, including scans, credit notes, missing purchase-order numbers, and awkward tax layouts. That number is a starting sample size, not a benchmark claim. Grade exact fields, record elapsed time in your own environment, and reject invalid output before it reaches an agent's screen. The architectural constraint follows from that test. Prompt text, response parsing, and business validation should not know which provider answered. Only configuration should. This keeps weekly shipping realistic: the team can change a model selection without rebuilding an adapter and retesting unrelated application code. Test the contract. There is a cost. Normalization targets the shared surface. Claude's, Gemini's, or OpenAI's newest native feature may appear before a multi-model layer exposes it. If your extraction quality depends on one such feature, use that provider directly and accept the coupling. Portability is a choice, not a free bonus. The script below discovers an available chat model, sends one extraction request through an OpenAI-compatible client, validates the answer, and reports elapsed time. It uses an environment variable for the key and an optional MODEL ID override, so model selection stays out of source control. Install openai and zod , then run it with a recent TypeScript runner. The sample input is text because document ingestion is a separate decision. OCR quality can dominate model quality; mixing both into the first comparison makes the result hard to interpret. python import OpenAI from "openai"; import { z } from "zod"; const apiKey = process.env.INFRAI API KEY; if apiKey throw new Error "INFRAI API KEY is required" ; const baseURL = process.env.INFRAI BASE URL; if baseURL throw new Error "INFRAI BASE URL is required" ; const ModelList = z.object { object: z.literal "list" , capability: z.string , available only: z.boolean , count: z.number , data: z.array z.object { id: z.string , capability: z.string , available: z.boolean , } .passthrough , , } ; const Invoice = z.object { invoice number: z.string , invoice date: z.string , currency: z.string .length 3 , total: z.number .nonnegative , purchase order: z.string .nullable , line items: z.array z.object { description: z.string , quantity: z.number .positive , unit price: z.number .nonnegative , } , } ; const sleep = ms: number = new Promise resolve = setTimeout resolve, ms ; async function fetchModels attempt = 0 : Promise