# Claix: Evidence-Backed Document JSON API for Agents

> Source: <https://idp-software.com/vendors/claix/>
> Published: 2026-09-29 16:58:10+00:00

# Claix: Evidence-Backed Document JSON API for Agents

A Madrid developer API that converts PDFs, spreadsheets, Word files, images, text and audio into JSON that follows a customer-defined schema, attaches a source location to every value, and keeps processed documents queryable for AI agents.

## Overview

Claix sells [data extraction](../../capabilities/data-extraction/) as a metered API for developers who build agents and automations rather than for operations teams who review documents. A file goes in through one of six format-specific endpoints, a JSON schema defines the output, and each field comes back as a value paired with where it was found: a page and fragment in a PDF, a column and row in a spreadsheet. Processed documents can be kept and queried later by ID, alone or grouped in a "Knowledge Space" for questions that span several files.

The business is very young and very small. The [terms of service](https://www.claix.dev/terms) name the operator as Gael Anaya Carballo, trading as Claix, a sole trader in Madrid working under Spanish law. The GitHub organization was opened in August 2026, the MCP server repository in September 2026, and the blog, which runs back to July 2026, is written by the founder. There's no announced funding, no named customer and no published accuracy figure. What Claix does publish is unusually specific for its size: a flat price per call, a full subprocessor list and a vendor-written comparison that says plainly where [Reducto](../../vendors/reducto-ai/) is the better choice.

The buyer is a developer or small team wiring documents into an agent, an n8n or Make workflow, or a backend service who wants to avoid running OCR, chunking and a vector database. Comparable APIs in this [IDP vendor directory](../../vendors/) include [Reducto](../../vendors/reducto-ai/), [LlamaParse](../../vendors/llamaparse/), [Unstructured](../../vendors/unstructured/), [Parseur](../../vendors/parseur/) and the Madrid platform [anyformat](../../vendors/anyformat/), which also returns field-level citations but targets enterprise workflows.

## Source tracing instead of a confidence score

The main design choice is how Claix signals doubt. Many extraction APIs attach a confidence score to each field. Claix [argues against that](https://www.claix.dev/blog/evidence-backed-document-intelligence-for-ai-agents): a score of 0.98 says nothing about where the number came from. With source verification switched on, every field is returned as a `{ value, source }` pair:

| **Input** | **What the source points to** | 
|---|---|
| Spreadsheet | Column name and row number | 
| PDF, Word, image | Page, paragraph, clause, table or quoted fragment | 
| Knowledge Space query | Document IDs, file names and the cross-document reasoning path | 

When the model can't locate a value, Claix doesn't invent a citation. The source is set to `requires_human_revision` and the value is left null, so a workflow can route that field to a person instead of passing on a plausible guess. That is a sound pattern for [quality verification](../../capabilities/quality-verification/) in agent pipelines, and it is easier to audit than a probability. The limits are stated by the vendor as well: a citation makes a value checkable, not correct, and Claix publishes no measurement of how often values or citations are wrong.

## Six endpoints and three ways in

Each format has its own endpoint under `https://claix.dev/api`: `excel-json` for XLSX and CSV, `pdf-json` for native and scanned PDFs, `doc-json` for Word, `img-json` for PNG, JPEG and WebP, `txt-json` for text, HTML, Markdown and XML, and `audio-json` for MP3, WAV, M4A and OGG recordings. Requests are multipart uploads with a `schema_id`, authenticated with an API key in an `x-api-key` or Bearer header. The [Excel documentation](https://www.claix.dev/documentation/excel-to-json) is candid about small constraints, such as only the first sheet of a workbook being processed, and lists which error codes are safe to retry. An OpenAPI file is published.

Beyond REST, Claix ships an [MCP server](https://github.com/claix-dev/claix-mcp) so Claude Desktop, Cursor, Windsurf and other MCP clients can extract and query documents as tools, and an A2A endpoint with a public Agent Card at `/.well-known/agent.json` so other agents can discover and delegate to it. An "agent mode" accepts a natural-language instruction instead of a fixed schema. For low-code users, the blog walks through calling the API from [n8n](https://www.claix.dev/blog/integrar-claix-en-n8n) and Make. This puts Claix firmly in the [agentic](../../capabilities/agentic/) end of the market: the documentation assumes the caller is software.

## Knowledge Spaces: document memory without a vector database

Documents can be processed and discarded, kept temporarily, or kept persistently. Persisted documents get a `document_id` and can be queried later with up to five questions per call, at €0.03, without being sent again. Grouping documents under a `space_id` lets a query compare, sum or reconcile across files, for example a supplier contract, its invoices and purchase orders. [Documents can be added to, removed from or replaced in a space](https://www.claix.dev/blog/dynamic-knowledge-spaces-ai) without changing their ID, and those management calls are free.

Claix pitches this as an alternative to stuffing whole PDFs into a prompt or building a retrieval stack. It doesn't document how retrieval works inside a space, how large a space can grow, or how answers are affected as the number of documents rises. Teams planning to put hundreds of files in one space should test that before relying on it.

## Pricing: flat per call, or bring your own key

Pricing is the most transparent part of the offer. Signup includes 100 successful calls with no credit card. After that, [each successful call costs €0.15](https://www.claix.dev/) regardless of file size or page count, and a context query costs €0.03. Failed calls, whether authentication errors, validation errors or server errors, are not billed. Under bring-your-own-key (BYOK), the customer supplies an OpenAI, Gemini, Anthropic or xAI key, pays that provider's token costs directly, and Claix charges no processing fee.

A flat per-call price favors long documents: a 60-page contract costs the same as a one-page receipt. Buyers comparing against per-page APIs should check two things. The pricing figures in older Claix blog posts, such as €0.10 to €0.15 per document in the Reducto comparison, differ from the current homepage, so the live pricing section is the one to rely on. And the BYOK route, being free, has no stated service level; how Claix funds that tier over time isn't explained.

## Data handling: EU hosting, US model providers

The homepage states data is hosted in Stockholm, with encryption in transit and at rest, API-key authentication and tenant isolation. Customers choose retention per document: instant deletion, temporary, or persistent. Claix says neither it nor its providers train on customer documents.

The [data processing agreement](https://www.claix.dev/dpa) fills in the rest. Managed AI runs on Google Gemini, with OpenAI and Anthropic listed as contingency providers; Supabase provides the database and authentication in the EEA, AWS provides cloud infrastructure depending on region, and internal workflows run on a self-hosted n8n instance in Frankfurt. Transfers outside the EEA rely on standard contractual clauses and the EU-US Data Privacy Framework. For [security and compliance](../../capabilities/security-compliance/) reviews, that means document content can reach US model providers even though storage sits in the EU, and it means the processor is a single individual rather than a company. No ISO 27001 or SOC 2 report is claimed, and there's no on-premises or VPC option. Claix's own Reducto comparison says so directly.

## Use cases

Claix publishes no customer case studies, so the use cases below come from its documentation and blog examples rather than from named deployments.

### Invoice and receipt extraction in automations

The homepage example is an invoice returned with number, supplier, taxable base and total, each with its source. The typical setup is an n8n or Make workflow that sends each incoming file to the matching endpoint and writes the JSON to a database, CRM or spreadsheet, with fields marked `requires_human_revision` routed to an approval step.

### Spreadsheet normalization

The Excel endpoint maps columns from arbitrary supplier or customer workbooks onto a fixed schema, recognizing synonyms, abbreviations and translated headers. This is where Claix started: its first blog post, in July 2026, is about Excel-to-JSON mapping.

### Agent memory over contracts and supporting documents

An agent processes a set of related files once, stores them in a Knowledge Space, and answers later questions, such as whether invoiced amounts exceed a contract ceiling, without resending the documents or holding them in its context window.

## Technical specifications

| **Feature** | **Specification** | 
|---|---|
| Input formats | PDF (native and scanned), XLSX, CSV, DOC, DOCX, PNG, JPEG, WebP, TXT, HTML, Markdown, XML, MP3, WAV, M4A, OGG | 
| Output | JSON validated against a customer schema; optional value and source pair per field | 
| Missing evidence | Value null, source `requires_human_revision` | 
| Document memory | Per-document queries by `document_id` ; cross-document queries by`space_id` | 
| Access | REST API with OpenAPI spec, MCP server, A2A endpoint with Agent Card | 
| Authentication | API key in `x-api-key` or Bearer header | 
| Models | Google Gemini by default; OpenAI and Anthropic as contingency; BYOK with OpenAI, Gemini, Anthropic or xAI | 
| Hosting | Stockholm per homepage; Supabase in the EEA, AWS, self-hosted n8n in Frankfurt | 
| Retention | Instant deletion, temporary or persistent, chosen by the customer | 
| Deployment | Managed cloud only; no on-premises or private cloud | 
| Certifications | None claimed; GDPR processor agreement | 
| Pricing | 100 successful calls free; €0.15 per call; €0.03 per context query; BYOK without Claix fee | 

## Resources

- [Claix website](https://www.claix.dev/)
- [API documentation](https://www.claix.dev/documentation/excel-to-json)
- [MCP server on GitHub](https://github.com/claix-dev/claix-mcp) and[BYOK examples](https://github.com/claix-dev/free-document-processing-with-byok)
- [Data processing agreement](https://www.claix.dev/dpa) and[terms](https://www.claix.dev/terms)
- [Blog](https://www.claix.dev/blog)

## Company information

Claix is operated by Gael Anaya Carballo as a sole trader

Madrid, Spain

info@claix.dev

claix.dev

Because the contracting party is an individual rather than a company, buyers should look closely at the liability, continuity and data-return terms before putting Claix into a production workflow, and keep their schemas and prompts portable. The BYOK option and the OpenAPI spec make switching away easier than with most closed APIs.
