{"slug": "structured-outputs-for-ai-generated-financial-models-schemas-before-spreadsheets", "title": "Structured Outputs for AI-Generated Financial Models: Schemas Before Spreadsheets", "summary": "A developer proposes representing AI-generated financial models as structured data validated against JSON Schema before rendering them into spreadsheets, arguing that LLM spreadsheet generation is fundamentally a data-contract problem. The approach encodes inputs, assumptions, units, and calculation formulas in an intermediate JSON representation so that structural and unit consistency can be checked deterministically before Excel is involved, using schema-constrained structured outputs from APIs such as OpenAI's. The author notes schema validation can confirm a value is numeric but cannot confirm the assumption itself is correct for a given project.", "body_md": "Generating a financial model with an LLM is not primarily a spreadsheet-generation problem. It is a data-contract problem. When an LLM is asked to create a financial model directly in Excel, several important decisions can become implicit. What is an input? Which numbers are assumptions? Which values came from external evidence? Which numbers are calculated? What units are being used? Which fields are mandatory?\n\nA spreadsheet can represent all of these things, but it is not necessarily the best place to define the contract between an LLM and a financial-modeling system. A more reliable approach is to first represent the model as structured data, validate that structure, perform deterministic calculations, and only then render the result into a spreadsheet.\n\nConsider a simple development model with inputs such as gross floor area, saleable area, selling price, construction cost, professional fees, financing assumptions, and development timing.\n\nA prompt such as:\n\nCreate an Excel development feasibility model from these assumptions. may produce a workbook that looks reasonable. That does not mean the model is correct.\n\nFor example, construction cost might be calculated using gross floor area:\n\nconstruction_cost = GFA × construction_cost_per_sqm\n\nBut another implementation could use saleable area:\n\nconstruction_cost = saleable_area × construction_cost_per_sqm\n\nBoth can result in valid Excel formulas.\n\nThe difference is a modelling decision, not a spreadsheet-formatting decision. The same problem occurs with percentages, currencies, time periods, unit conversions, missing assumptions, and the distinction between user inputs and derived values.\n\nThe system needs to represent those decisions explicitly before Excel becomes involved.\n\nA useful intermediate representation can describe the financial model without depending on a particular spreadsheet.\n\nFor example:\n\n{\n\n  \"model\": {\n\n    \"name\": \"development_feasibility\",\n\n    \"currency\": \"USD\",\n\n    \"period\": \"monthly\"\n\n  },\n\n  \"inputs\": [\n\n    {\n\n      \"id\": \"gfa\",\n\n      \"value\": 25000,\n\n      \"unit\": \"sqm\",\n\n      \"source_type\": \"user_input\"\n\n    },\n\n    {\n\n      \"id\": \"construction_cost\",\n\n      \"value\": 1800,\n\n      \"unit\": \"USD/sqm\",\n\n      \"source_type\": \"assumption\"\n\n    }\n\n  ],\n\n  \"calculations\": [\n\n    {\n\n      \"id\": \"construction_cost_total\",\n\n      \"formula\": \"gfa * construction_cost\",\n\n      \"unit\": \"USD\"\n\n    }\n\n  ]\n\n}\n\nThe representation is deliberately simple. The important part is that the model's meaning exists independently of the workbook. The application can validate the structure before any spreadsheet is generated. It can also determine whether a calculation references a known input and whether the units and expected data types are consistent.\n\nJSON Schema is useful because it defines the expected structure of the data. A schema can require that value is a number, that unit is a string, and that source_type belongs to an approved set of values.\n\nFor example:\n\n{\n\n  \"type\": \"object\",\n\n  \"properties\": {\n\n    \"id\": {\n\n      \"type\": \"string\"\n\n    },\n\n    \"value\": {\n\n      \"type\": \"number\"\n\n    },\n\n    \"unit\": {\n\n      \"type\": \"string\"\n\n    },\n\n    \"source_type\": {\n\n      \"type\": \"string\",\n\n      \"enum\": [\n\n        \"user_input\",\n\n        \"assumption\",\n\n        \"verified_evidence\"\n\n      ]\n\n    }\n\n  },\n\n  \"required\": [\n\n    \"id\",\n\n    \"value\",\n\n    \"unit\",\n\n    \"source_type\"\n\n  ],\n\n  \"additionalProperties\": false\n\n}\n\nThis prevents the model from returning something structurally invalid, such as a text description where a numeric value is required. Modern structured-output APIs can use JSON Schema to constrain model responses. OpenAI's current documentation also distinguishes schema-constrained structured outputs from simply requesting valid JSON. But schema validation has a clear limit.\n\nIt can establish that:\n\n{\n\n  \"construction_cost\": 1800\n\n}\n\ncontains a number. It cannot establish that 1800 is the correct construction-cost assumption for a particular project. That requires evidence or professional judgement. This distinction is important when designing systems for financial modelling.\n\nThere should be more than one validation layer. The first layer checks the structure of the response. It can validate required fields, data types, enumerations, object structure, and permitted properties.\n\nThe second layer checks whether the proposed model makes sense. It can check units, dependencies, calculation references, period consistency, missing assumptions, sign conventions, and other modelling rules.\n\nA response can pass the first layer and still fail the second. For example, this is structurally valid:\n\n{\n\n  \"gfa\": 25000,\n\n  \"construction_cost\": 1800,\n\n  \"currency\": \"USD\"\n\n}\n\nBut if the calculation engine expects construction cost in USD per square foot, the model has a semantic problem even though every field has the correct basic data type. A schema therefore should not be treated as a financial-model validator. It is one part of the validation system.\n\nLLMs are useful for interpreting natural-language instructions and converting them into structured representations. They are less appropriate as the final authority for calculations that can be expressed deterministically.\n\nSuppose the model specification contains:\n\ngfa = 25,000 sqm\n\nconstruction_cost = 1,800 USD/sqm\n\nThe calculation:\n\n25,000 × 1,800\n\ndoes not need probabilistic reasoning.\n\nThe calculation engine should perform it. This separation makes the system easier to test because the calculation can be evaluated independently of the language model. It also makes changes easier to trace. If the construction-cost assumption changes, the system can identify the affected calculation rather than relying on the LLM to regenerate an entire workbook.\n\nExcel remains useful because financial professionals can inspect formulas, modify assumptions, test scenarios, and work with a familiar interface. The important design decision is not to remove Excel. It is to avoid making Excel the only representation of the model.\n\nThe Excel JavaScript API provides ranges for reading and writing values and formulas. Microsoft's documentation describes Range as the basic object for working with cells and contiguous blocks, with separate values and formulas properties. That means a validated model specification can be translated into a workbook in a controlled way.\n\nFor example, the renderer might write an input value to a specific range and write a formula to another range.\n\nconst inputRange = worksheet.getRange(\"B5\");\n\ninputRange.values = [[25000]];\n\nconst outputRange = worksheet.getRange(\"B10\");\n\noutputRange.formulas = [[\"=B5*B6\"]];\n\nThe important point is that the formula is being generated from an already validated model definition. The LLM is not deciding what every spreadsheet cell should contain during workbook construction.\n\nFinancial models contain different kinds of numbers. A user-provided input is not the same thing as an external market observation. An assumption is not the same thing as a calculated output. Those distinctions should remain visible in the structured representation.\n\nFor example:\n\n{\n\n  \"id\": \"sale_price\",\n\n  \"value\": 4200,\n\n  \"unit\": \"USD/sqm\",\n\n  \"source_type\": \"verified_evidence\",\n\n  \"source_id\": \"market_source_017\"\n\n}\n\nis materially different from:\n\n{\n\n  \"id\": \"sale_price\",\n\n  \"value\": 4200,\n\n  \"unit\": \"USD/sqm\",\n\n  \"source_type\": \"assumption\"\n\n}\n\nThe numeric value is identical. The provenance is not. A system should also avoid allowing the LLM to invent source identifiers. If a source identifier represents external evidence, it should come from the evidence layer or from an explicitly supplied source. Otherwise the system can create the appearance of provenance without actually having verifiable provenance.\n\nThere is a temptation to put everything into the schema.\n\nEvery cell.\n\nEvery formula.\n\nEvery formatting property.\n\nEvery Excel address.\n\nEvery chart.\n\nThat creates another problem.\n\nThe schema should represent the meaning of the financial model rather than reproduce the entire workbook.\n\nThis is useful:\n\nrevenue = units × price_per_unit\n\nThis is implementation detail:\n\ncell = H47\n\nformula = \"=SUM(H31:H46)\"\n\nfont = \"Calibri\"\n\nfont_size = 11\n\nThe first describes model logic. The second describes spreadsheet presentation. Keeping those concerns separate makes the system easier to maintain.\n\nThe model schema will evolve.\n\nA first version might contain:\n\ngfa\n\nconstruction_cost\n\nselling_price\n\nA later version might introduce:\n\ngfa\n\nsaleable_area\n\nconstruction_cost\n\nselling_price\n\nphasing\n\nIf older model specifications cannot be identified or interpreted after a schema change, reproducibility becomes difficult.\n\nA simple version field can help:\n\n{\n\n  \"schema_version\": \"1.2\"\n\n}\n\nSchema changes can then be tested in the same way that API changes are tested. Maintain representative model fixtures and validate them whenever the schema or calculation engine changes. This is particularly important when an LLM is involved because the model may produce different structures as prompts, schemas, or model versions change.\n\nStructured outputs do not make the underlying financial assumptions correct.\n\nThat boundary makes the system easier to validate, test, debug, and audit.\n\nFor AI-generated financial models, the useful principle is simple. Let the LLM interpret the request and propose a structured representation. Validate that representation before using it. Keep material calculations deterministic. Preserve the provenance of important inputs. Use Excel to expose the resulting model to the professional user. The interesting engineering question is therefore not whether an LLM can create an Excel file.\n\nIt can.\n\nThe more important question is whether the system can explain what each material number represents, where the number came from, how it was calculated, and which assumptions affect it. That is why structured outputs should come before spreadsheets. AI-Assisted. Human technical review required before publication.", "url": "https://wpnews.pro/news/structured-outputs-for-ai-generated-financial-models-schemas-before-spreadsheets", "canonical_source": "https://dev.to/feasibilityproaiai/structured-outputs-for-ai-generated-financial-models-schemas-before-spreadsheets-3e0h", "published_at": "2026-10-03 06:04:37+00:00", "updated_at": "2026-10-03 06:07:54.746844+00:00", "lang": "en", "topics": ["structured-data", "large-language-models", "ai-tools", "developer-tools"], "entities": ["OpenAI", "JSON Schema", "Excel"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/structured-outputs-for-ai-generated-financial-models-schemas-before-spreadsheets", "markdown": "https://wpnews.pro/news/structured-outputs-for-ai-generated-financial-models-schemas-before-spreadsheets.md", "text": "https://wpnews.pro/news/structured-outputs-for-ai-generated-financial-models-schemas-before-spreadsheets.txt", "jsonld": "https://wpnews.pro/news/structured-outputs-for-ai-generated-financial-models-schemas-before-spreadsheets.jsonld"}}