What if Jev spoke Arrow? TypeSafe AI's new model Jev turns natural language and application state into typed decisions, returning choices, scores, and probabilities as JSON that code can consume directly, and the company reports substantial speed and cost gains over general-purpose LLMs in decision workflows. TypeSafe describes Jev as built on a new model architecture, a parallel sampler, and a training method called Reinforcement Learning for Calibrated Decisions, producing probabilities in parallel rather than generating answers token by token. The post explores modeling Jev's three question types — Choice, Noul, and Score — as Apache Arrow columns, using Arrow extension types to attach semantic meaning to the storage types. What if Jev spoke Arrow? Jev https://docs.typesafe.ai/concepts/system-one is TypeSafe AI’s new model for turning natural language and application state into typed decisions. You supply the context and define the possible answers. Jev returns choices, scores, and probabilities that your code can use directly. The API https://docs.typesafe.ai/api delivers those answers as JSON. If you’ve managed to avoid hearing about Jev lately, the rock you’re hiding under has excellent soundproofing. Jev is part of a wider effort to make AI outputs easier to use in code. Other tools, such as Outlines from .txt https://dottxt.ai/ , use constrained decoding to make existing language models produce outputs that conform to a schema. TypeSafe took a different approach. In its announcement https://typesafe.ai/blog/introducing-system-one-models-and-jev , the company describes a new model architecture, a parallel sampler, and a training method called Reinforcement Learning for Calibrated Decisions. Jev produces probabilities in parallel, avoiding the work of generating an answer token by token. TypeSafe reports substantial gains in speed and cost compared with general-purpose LLMs in its decision workflows. Jev is a new primitive, and nobody yet knows the full scope of what it will make possible. This post reflects our thinking at this point in time, and we expect it to evolve. But some powerful patterns are already clear. TypeSafe’s docs describe several https://docs.typesafe.ai/patterns . Speculative fan-out https://docs.typesafe.ai/patterns/fan-out asks many questions in one call, including speculative ones, and lets your code decide which answers are relevant. Confidence-gated routing https://docs.typesafe.ai/patterns/confidence-routing treats confidence as a second decision axis, so your code can take a different path when Jev is unsure. Composite scoring https://docs.typesafe.ai/patterns/composite-scoring combines several dimensions of judgment into a single score. Intent routing https://docs.typesafe.ai/patterns/intent-routing classifies what a user wants and sends the request to the right handler. Together, these patterns open the door to fundamentally probabilistic workflows and pipelines, in places where until recently we would have assumed only deterministic ones were practical. The sophistication you can achieve is astounding. Pipelines like these move a lot of structured data, which got us curious about how Jev might work with another technology that combines structure with performance and efficiency: Apache Arrow https://arrow.apache.org/ . In Stop paying the JSON tax https://columnar.tech/blog/stop-paying-the-json-tax/ , we described how Arrow can speed up data pipelines by avoiding conversions to JSON and back. Could it do that here? Modeling Jev answers as Arrow To explore that question, we first designed an Arrow schema to represent Jev’s answers. Jev has three question types https://docs.typesafe.ai/primitives , each with a different answer shape: | Question type | What it returns | |---|---| | Choice | A selected option, a probability for every option, and a confidence score. | | Noul | The probability that the answer to a yes/no question is yes. There is no separate confidence field. | | Score | A position on ordered levels, which can fall between levels; a probability for each level, confidence, and a legend describing the levels. | We chose to represent each question’s answers as an Arrow column. The question’s definition tells us the column’s type before inference starts. Choice labels and Score legends describe the possible answers, so we can put them in the schema’s field metadata. The predictions and probabilities go in the data buffers. Arrow extension types let us attach that semantic meaning to ordinary Arrow storage types: Choice struct< choice: uint8 not null, confidence: float64 not null, probabilities: fixed size list