Claix: Evidence-Backed Document JSON API for Agents Madrid sole trader Gael Anaya Carballo, trading as Claix, has launched a metered document-extraction API that returns every extracted field as a { value, source } pair pointing to a page, paragraph, clause, table or quoted fragment, or to a column and row in a spreadsheet, instead of a confidence score. Claix exposes six format-specific endpoints under https://claix.dev/api — excel-json, pdf-json, doc-json, img-json, txt-json and audio-json — authenticated with an API key in an x-api-key or Bearer header, and keeps processed documents queryable by ID or grouped in a Knowledge Space. When the model cannot locate a value, Claix sets the source to requires_human_revision and leaves the value null rather than inventing a citation; the company publishes no funding, named customer or accuracy figure. Claix: Evidence-Backed Document JSON API for Agents A Madrid developer API that converts PDFs, spreadsheets, Word files, images, text and audio into JSON that follows a customer-defined schema, attaches a source location to every value, and keeps processed documents queryable for AI agents. Overview Claix sells data extraction ../../capabilities/data-extraction/ as a metered API for developers who build agents and automations rather than for operations teams who review documents. A file goes in through one of six format-specific endpoints, a JSON schema defines the output, and each field comes back as a value paired with where it was found: a page and fragment in a PDF, a column and row in a spreadsheet. Processed documents can be kept and queried later by ID, alone or grouped in a "Knowledge Space" for questions that span several files. The business is very young and very small. The terms of service https://www.claix.dev/terms name the operator as Gael Anaya Carballo, trading as Claix, a sole trader in Madrid working under Spanish law. The GitHub organization was opened in August 2026, the MCP server repository in September 2026, and the blog, which runs back to July 2026, is written by the founder. There's no announced funding, no named customer and no published accuracy figure. What Claix does publish is unusually specific for its size: a flat price per call, a full subprocessor list and a vendor-written comparison that says plainly where Reducto ../../vendors/reducto-ai/ is the better choice. The buyer is a developer or small team wiring documents into an agent, an n8n or Make workflow, or a backend service who wants to avoid running OCR, chunking and a vector database. Comparable APIs in this IDP vendor directory ../../vendors/ include Reducto ../../vendors/reducto-ai/ , LlamaParse ../../vendors/llamaparse/ , Unstructured ../../vendors/unstructured/ , Parseur ../../vendors/parseur/ and the Madrid platform anyformat ../../vendors/anyformat/ , which also returns field-level citations but targets enterprise workflows. Source tracing instead of a confidence score The main design choice is how Claix signals doubt. Many extraction APIs attach a confidence score to each field. Claix argues against that https://www.claix.dev/blog/evidence-backed-document-intelligence-for-ai-agents : a score of 0.98 says nothing about where the number came from. With source verification switched on, every field is returned as a { value, source } pair: | Input | What the source points to | |---|---| | Spreadsheet | Column name and row number | | PDF, Word, image | Page, paragraph, clause, table or quoted fragment | | Knowledge Space query | Document IDs, file names and the cross-document reasoning path | When the model can't locate a value, Claix doesn't invent a citation. The source is set to requires human revision and the value is left null, so a workflow can route that field to a person instead of passing on a plausible guess. That is a sound pattern for quality verification ../../capabilities/quality-verification/ in agent pipelines, and it is easier to audit than a probability. The limits are stated by the vendor as well: a citation makes a value checkable, not correct, and Claix publishes no measurement of how often values or citations are wrong. Six endpoints and three ways in Each format has its own endpoint under https://claix.dev/api : excel-json for XLSX and CSV, pdf-json for native and scanned PDFs, doc-json for Word, img-json for PNG, JPEG and WebP, txt-json for text, HTML, Markdown and XML, and audio-json for MP3, WAV, M4A and OGG recordings. Requests are multipart uploads with a schema id , authenticated with an API key in an x-api-key or Bearer header. The Excel documentation https://www.claix.dev/documentation/excel-to-json is candid about small constraints, such as only the first sheet of a workbook being processed, and lists which error codes are safe to retry. An OpenAPI file is published. Beyond REST, Claix ships an MCP server https://github.com/claix-dev/claix-mcp so Claude Desktop, Cursor, Windsurf and other MCP clients can extract and query documents as tools, and an A2A endpoint with a public Agent Card at /.well-known/agent.json so other agents can discover and delegate to it. An "agent mode" accepts a natural-language instruction instead of a fixed schema. For low-code users, the blog walks through calling the API from n8n https://www.claix.dev/blog/integrar-claix-en-n8n and Make. This puts Claix firmly in the agentic ../../capabilities/agentic/ end of the market: the documentation assumes the caller is software. Knowledge Spaces: document memory without a vector database Documents can be processed and discarded, kept temporarily, or kept persistently. Persisted documents get a document id and can be queried later with up to five questions per call, at €0.03, without being sent again. Grouping documents under a space id lets a query compare, sum or reconcile across files, for example a supplier contract, its invoices and purchase orders. Documents can be added to, removed from or replaced in a space https://www.claix.dev/blog/dynamic-knowledge-spaces-ai without changing their ID, and those management calls are free. Claix pitches this as an alternative to stuffing whole PDFs into a prompt or building a retrieval stack. It doesn't document how retrieval works inside a space, how large a space can grow, or how answers are affected as the number of documents rises. Teams planning to put hundreds of files in one space should test that before relying on it. Pricing: flat per call, or bring your own key Pricing is the most transparent part of the offer. Signup includes 100 successful calls with no credit card. After that, each successful call costs €0.15 https://www.claix.dev/ regardless of file size or page count, and a context query costs €0.03. Failed calls, whether authentication errors, validation errors or server errors, are not billed. Under bring-your-own-key BYOK , the customer supplies an OpenAI, Gemini, Anthropic or xAI key, pays that provider's token costs directly, and Claix charges no processing fee. A flat per-call price favors long documents: a 60-page contract costs the same as a one-page receipt. Buyers comparing against per-page APIs should check two things. The pricing figures in older Claix blog posts, such as €0.10 to €0.15 per document in the Reducto comparison, differ from the current homepage, so the live pricing section is the one to rely on. And the BYOK route, being free, has no stated service level; how Claix funds that tier over time isn't explained. Data handling: EU hosting, US model providers The homepage states data is hosted in Stockholm, with encryption in transit and at rest, API-key authentication and tenant isolation. Customers choose retention per document: instant deletion, temporary, or persistent. Claix says neither it nor its providers train on customer documents. The data processing agreement https://www.claix.dev/dpa fills in the rest. Managed AI runs on Google Gemini, with OpenAI and Anthropic listed as contingency providers; Supabase provides the database and authentication in the EEA, AWS provides cloud infrastructure depending on region, and internal workflows run on a self-hosted n8n instance in Frankfurt. Transfers outside the EEA rely on standard contractual clauses and the EU-US Data Privacy Framework. For security and compliance ../../capabilities/security-compliance/ reviews, that means document content can reach US model providers even though storage sits in the EU, and it means the processor is a single individual rather than a company. No ISO 27001 or SOC 2 report is claimed, and there's no on-premises or VPC option. Claix's own Reducto comparison says so directly. Use cases Claix publishes no customer case studies, so the use cases below come from its documentation and blog examples rather than from named deployments. Invoice and receipt extraction in automations The homepage example is an invoice returned with number, supplier, taxable base and total, each with its source. The typical setup is an n8n or Make workflow that sends each incoming file to the matching endpoint and writes the JSON to a database, CRM or spreadsheet, with fields marked requires human revision routed to an approval step. Spreadsheet normalization The Excel endpoint maps columns from arbitrary supplier or customer workbooks onto a fixed schema, recognizing synonyms, abbreviations and translated headers. This is where Claix started: its first blog post, in July 2026, is about Excel-to-JSON mapping. Agent memory over contracts and supporting documents An agent processes a set of related files once, stores them in a Knowledge Space, and answers later questions, such as whether invoiced amounts exceed a contract ceiling, without resending the documents or holding them in its context window. Technical specifications | Feature | Specification | |---|---| | Input formats | PDF native and scanned , XLSX, CSV, DOC, DOCX, PNG, JPEG, WebP, TXT, HTML, Markdown, XML, MP3, WAV, M4A, OGG | | Output | JSON validated against a customer schema; optional value and source pair per field | | Missing evidence | Value null, source requires human revision | | Document memory | Per-document queries by document id ; cross-document queries by space id | | Access | REST API with OpenAPI spec, MCP server, A2A endpoint with Agent Card | | Authentication | API key in x-api-key or Bearer header | | Models | Google Gemini by default; OpenAI and Anthropic as contingency; BYOK with OpenAI, Gemini, Anthropic or xAI | | Hosting | Stockholm per homepage; Supabase in the EEA, AWS, self-hosted n8n in Frankfurt | | Retention | Instant deletion, temporary or persistent, chosen by the customer | | Deployment | Managed cloud only; no on-premises or private cloud | | Certifications | None claimed; GDPR processor agreement | | Pricing | 100 successful calls free; €0.15 per call; €0.03 per context query; BYOK without Claix fee | Resources - Claix website https://www.claix.dev/ - API documentation https://www.claix.dev/documentation/excel-to-json - MCP server on GitHub https://github.com/claix-dev/claix-mcp and BYOK examples https://github.com/claix-dev/free-document-processing-with-byok - Data processing agreement https://www.claix.dev/dpa and terms https://www.claix.dev/terms - Blog https://www.claix.dev/blog Company information Claix is operated by Gael Anaya Carballo as a sole trader Madrid, Spain info@claix.dev claix.dev Because the contracting party is an individual rather than a company, buyers should look closely at the liability, continuity and data-return terms before putting Claix into a production workflow, and keep their schemas and prompts portable. The BYOK option and the OpenAPI spec make switching away easier than with most closed APIs.