For supplier invoice extraction, an in-app chatbot should put OpenAI, Claude, and every other AI API behind one application-owned adapter. Run one recorded conversation unchanged against every candidate, then select on accepted JSON extractions, not token pricing or an advertised context window.
| Architecture | Portability | Maintenance | Use it when |
|---|---|---|---|
| Thin application adapter | High | Moderate | One SaaS owns the workflow |
| Self-hosted gateway behind an adapter | High | Higher | Several apps need shared policy |
| Provider-specific integration | Low | Low at first | One verified capability is mandatory |
Recommendation: keep the invoice schema and validation in the application. Put model IDs, limits, and request-envelope differences inside adapters. This spends a solo founder's scarce hours on review UX and extraction quality while leaving inference as a replaceable dependency.
The model should not define the invoice record. The application should. A useful record might contain supplierName, invoiceNumber, currency, and totalMinor. A chat turn may clarify a missing currency, but it must not change field names because an API uses a different structured-output envelope.
Keep the conversation, extraction request, and normalized result separate. Conversation messages are user-facing history. The request contains only evidence needed for this turn. The result is domain data validated before it reaches reconciliation or a database.
Small wins here.
type Message = {
role: "system" | "user" | "assistant";
content: string;
};
type InvoiceExtraction = {
supplierName: string | null;
invoiceNumber: string | null;
currency: string | null;
totalMinor: number | null;
};
type ExtractResult = {
data: InvoiceExtraction;
usage: { inputTokens: number; outputTokens: number } | null;
finishReason: "complete" | "length" | "blocked" | "unknown";
};
interface InvoiceModel {
extract(messages: Message[]): Promise<ExtractResult>;
}
The interface omits vendor error classes, raw response envelopes, and prices. Adapters map those details. Version the domain schema separately so a weekly release has a stable target. This boundary has a real limitation: it exposes only common behavior, so a unique provider feature either needs a narrow extension or remains unavailable. That trade-off is acceptable when provider portability matters more than immediate access to every feature. It is not suitable when one exclusive capability defines the product.
Valid JSON is not necessarily a valid invoice. JSON can contain a string where the workflow requires an integer number of minor currency units. It cannot decide whether an empty invoice number means null, or whether a displayed total includes tax. Those are domain rules.
Validate every adapter result. Reject unknown keys and floating-point money. Check currency against the application's accepted representation. If line items are extracted, verify their arithmetic using an explicit tolerance rather than silently rewriting the stated total.
Allow one repair attempt for invalid structure. Stop there. An unbounded retry loop hides quality regressions and changes latency and usage. After the budget is exhausted, create a typed review task that shows the source evidence and disputed fields.
OpenAI, Anthropic Claude, Google Gemini, and OpenRouter can all enter the same trial as candidates. Product limits, model identifiers, and rate cards change, so verify their current primary documentation during the trial instead of freezing those values into domain code. None gets a special expected result. The same redacted corpus and acceptance rules apply to each. Require two adapters to pass before calling the contract portable, and allow exactly one structure-repair attempt per extraction. Those numbers are small on purpose: the first catches assumptions hidden by a single implementation, while the second prevents repair loops from disguising an invalid response.
Context capacity is a ceiling, not a batching target. The assistant normally needs the current document, extraction instructions, and a short relevant correction trail. Old invoices can introduce the wrong number or currency.
Build each request from a budget. Reserve output space. Include current invoice text first, then only corrections relevant to missing fields. If the document exceeds the tested budget, route it to an explicit split-and-merge process or manual review. Arbitrary truncation is dangerous because supplier details and totals may sit at opposite ends.
Do not guess.
type AdapterConfig = {
model: string;
maxInputTokens: number;
reservedOutputTokens: number;
};
function assertBudget(inputTokens: number, config: AdapterConfig): void {
const usable = config.maxInputTokens - config.reservedOutputTokens;
if (inputTokens > usable) throw new Error("invoice_input_over_budget");
}
Calibrate estimation for each adapter. Store actual usage when the response reports it; otherwise retain null. Pricing belongs in the evaluation harness. Compare cost per accepted extraction, alongside latency and review rate. Cost per call rewards cheap responses that may still be rejected.
Start with replay fixtures: redacted invoice text, short message history, expected fields, and allowed alternatives for genuine ambiguity. Run deterministic validation in CI. Run external calls on a controlled schedule so network variance does not make normal builds flaky.
The production path stays plain: authenticate, load the invoice, build a bounded request, call an adapter with a timeout, validate, then persist an accepted result or review task. Log a request ID, configuration version, duration, reported usage, validation outcome, and retry count. Do not log raw invoice text by default; supplier documents may contain address, account, and payment data.
Server-Sent Events fit one-way server-to-browser chat updates. The browser receives an event stream, but partial text must never become a committed invoice record. Complete the structured response, validate it, and only then notify the UI that fields are ready for review.
Ship one adapter and the replay harness in a weekly slice. Add a second adapter before calling the boundary portable. Otherwise, the interface only reflects its first implementation.
The runner-up becomes useful under a different constraint.
A self-hosted gateway is useful when multiple applications need shared routing, credentials, budgets, or observability. LiteLLM is one open-source example. Keep the gateway behind the application interface because its fields should not spread into invoice code. The cost is operational ownership: upgrades, credentials, availability, and logging become part of the workload.
A direct provider-specific integration can be reasonable when a verified requirement depends on one capability and its value exceeds the switching cost. Record the exception, isolate it in an adapter, and state what evidence would trigger reevaluation.
Keep the domain contract stable and the model choice reversible. That supports weekly shipping without turning an undifferentiated API decision into the center of the product.