{"slug": "what-if-jev-spoke-arrow", "title": "What if Jev spoke Arrow?", "summary": "TypeSafe AI's new model Jev turns natural language and application state into typed decisions, returning choices, scores, and probabilities as JSON that code can consume directly, and the company reports substantial speed and cost gains over general-purpose LLMs in decision workflows. TypeSafe describes Jev as built on a new model architecture, a parallel sampler, and a training method called Reinforcement Learning for Calibrated Decisions, producing probabilities in parallel rather than generating answers token by token. The post explores modeling Jev's three question types — Choice, Noul, and Score — as Apache Arrow columns, using Arrow extension types to attach semantic meaning to the storage types.", "body_md": "# What if Jev spoke Arrow?\n\n[Jev](https://docs.typesafe.ai/concepts/system-one) is TypeSafe AI’s new model for turning natural language and application state into typed decisions. You supply the context and define the possible answers. Jev returns choices, scores, and probabilities that your code can use directly. The [API](https://docs.typesafe.ai/api) delivers those answers as JSON. If you’ve managed to avoid hearing about Jev lately, the rock you’re hiding under has excellent soundproofing.\n\nJev is part of a wider effort to make AI outputs easier to use in code. Other tools, such as Outlines from [.txt](https://dottxt.ai/), use constrained decoding to make existing language models produce outputs that conform to a schema. TypeSafe took a different approach. In its [announcement](https://typesafe.ai/blog/introducing-system-one-models-and-jev), the company describes a new model architecture, a parallel sampler, and a training method called Reinforcement Learning for Calibrated Decisions. Jev produces probabilities in parallel, avoiding the work of generating an answer token by token. TypeSafe reports substantial gains in speed and cost compared with general-purpose LLMs in its decision workflows.\n\nJev is a new primitive, and nobody yet knows the full scope of what it will make possible. This post reflects our thinking at this point in time, and we expect it to evolve. But some powerful patterns are already clear. TypeSafe’s docs describe [several](https://docs.typesafe.ai/patterns). [Speculative fan-out](https://docs.typesafe.ai/patterns/fan-out) asks many questions in one call, including speculative ones, and lets your code decide which answers are relevant. [Confidence-gated routing](https://docs.typesafe.ai/patterns/confidence-routing) treats confidence as a second decision axis, so your code can take a different path when Jev is unsure. [Composite scoring](https://docs.typesafe.ai/patterns/composite-scoring) combines several dimensions of judgment into a single score. [Intent routing](https://docs.typesafe.ai/patterns/intent-routing) classifies what a user wants and sends the request to the right handler.\n\nTogether, these patterns open the door to fundamentally probabilistic workflows and pipelines, in places where until recently we would have assumed only deterministic ones were practical. The sophistication you can achieve is astounding.\n\nPipelines like these move a lot of structured data, which got us curious about how Jev might work with another technology that combines structure with performance and efficiency: [Apache Arrow](https://arrow.apache.org/). In [Stop paying the JSON tax](https://columnar.tech/blog/stop-paying-the-json-tax/), we described how Arrow can speed up data pipelines by avoiding conversions to JSON and back. Could it do that here?\n\n## Modeling Jev answers as Arrow\n\nTo explore that question, we first designed an Arrow schema to represent Jev’s answers. Jev has [three question types](https://docs.typesafe.ai/primitives), each with a different answer shape:\n\n| Question type | What it returns | \n|---|---|\n| Choice | A selected option, a probability for every option, and a confidence score. | \n| Noul | The probability that the answer to a yes/no question is yes. There is no separate confidence field. | \n| Score | A position on ordered levels, which can fall between levels; a probability for each level, confidence, and a legend describing the levels. | \n\nWe chose to represent each question’s answers as an Arrow column. The question’s definition tells us the column’s type before inference starts. Choice labels and Score legends describe the possible answers, so we can put them in the schema’s field metadata. The predictions and probabilities go in the data buffers. Arrow extension types let us attach that semantic meaning to ordinary Arrow storage types:\n\n```\nChoice\n  struct<\n    choice: uint8 not null,\n    confidence: float64 not null,\n    probabilities: fixed_size_list<item: float64 not null>[N] not null\n  > not null\n  metadata: {\"labels\":[\"returns\",\"shipping\",\"billing\",\"other\"]}\n\nNoul\n  float64 not null\n  metadata: {}\n\nScore\n  struct<\n    score: float64 not null,\n    confidence: float64 not null,\n    probabilities: fixed_size_list<item: float64 not null>[N] not null\n  > not null\n  metadata: {\"legend\":[\"Can wait\",\"Within a few days\",\"Today\"]}\n```\n\n`N` is the number of choices or score levels. Choice stores a one-byte index into the shared labels. Fixed-size probability vectors need no per-row offsets, and 64-bit floats preserve the values returned by the TypeSafe Python SDK. The schema metadata supplies the meaning of our choice indices and score levels. Together, the values and metadata are enough to reconstruct each original answer object.\n\nThe TypeSafe API returns these answers as JSON today. If it returned Arrow directly using this schema, tools such as [pandas](https://pandas.pydata.org/docs/user_guide/pyarrow.html), [Polars](https://docs.pola.rs/api/python/stable/reference/api/polars.from_arrow.html), [DuckDB](https://duckdb.org/docs/current/guides/python/sql_on_arrow), and [Apache DataFusion](https://datafusion.apache.org/python/user-guide/io/arrow.html) could consume the results without first deserializing JSON and rebuilding typed columns. Compatible consumers can use the Arrow buffers without copying them.\n\nThis is already a common way to exchange structured data. [Databricks](https://docs.databricks.com/aws/en/dev-tools/sql-execution-tutorial), [Snowflake](https://www.snowflake.com/en/blog/fetching-query-results-from-snowflake-just-got-a-lot-faster-with-apache-arrow/), and [ClickHouse](https://clickhouse.com/docs/reference/formats/Arrow/ArrowStream) can return query results in Arrow format over HTTP. [Hugging Face Datasets](https://huggingface.co/docs/datasets/use_with_pyarrow) uses Arrow internally and lets you retrieve Arrow tables directly. Giving Jev an Arrow output option would let its answers join those same data workflows.\n\n## From one state to many\n\nJev doesn’t offer Arrow output today, so our goal was to simulate what using the TypeSafe API would look like if it did. That meant taking Jev from the operational layer, where it makes decisions one interaction at a time, to the analytic layer, where the unit of work is a whole table. There, we hit a more basic problem with the request format.\n\nA [TypeSafe API request](https://docs.typesafe.ai/api) contains a `state` (the context to evaluate) and `questions` (a map of the judgments to make about it). The API is wonderfully designed for asking multiple questions about *one* state. But it has no native batch operation for asking the same set of questions about *many* independent states.\n\nFor example, take a live customer interaction. You might want to identify intent, score urgency, check refund eligibility, and look for signs of fraud. You can send the context once, ask all those questions together, and Jev evaluates them independently in parallel. The [documentation makes good use of this](https://docs.typesafe.ai/cookbooks/parallel_questions). But before trusting those judgments in a live application, you’d want to validate them. Part of that is reasoning about the probabilities Jev should produce across the range of states it might see. But you’d also want to run it against a representative sample of historical customer interactions and compare its answers with known outcomes. That’s the bulk workload we’re interested in: lots of states, a fixed set of questions. Jev’s speed and pricing make it an appealing fit for that work. The obstacle is the API: it has no bulk endpoint, so you need to make thousands or millions of separate API calls, one per state.\n\n## Building Jevaro\n\nWe weren’t ready to give up on the experiment, so we built [Jevaro](https://github.com/columnar-tech/jevaro): a small Python proxy server, with Python and JavaScript clients. It accepts multiple `states` and a shared `questions` map in one request, calls Jev for each state, and returns an Arrow IPC stream using the [schema above](https://github.com/columnar-tech/jevaro/blob/main/docs/arrow-schema.md). Results arrive in input order. The schema goes out immediately, and answers follow as they become available in that order.\n\nImproving throughput took several rounds of tuning. Opening fresh connections repeatedly adds network and TLS setup time; waiting for each answer before sending the next request prevents calls from overlapping. We reused a long-lived HTTP/2 connection pool, kept a window of asynchronous requests pending, and refilled it before writing result batches. SDK retries recover dropped connections and back off on rate limits and overload responses. All of this preserves the input order of the results.\n\nEven after that tuning, every state still needs a separate upstream HTTP request and JSON response. We’d expect a native bulk call returning an Arrow stream directly from the API to achieve orders of magnitude better throughput by avoiding all that repeated request handling and JSON conversion. At a minimum, it could shift the bottleneck from API overhead to inference itself.\n\nFor now, Jevaro lets us experiment with that interface. We could have put the request logic directly in each SDK, but a batching proxy server gives us one implementation of concurrency, retries, ordering, and Arrow serialization. Browser clients can also use it without receiving the TypeSafe key. The cost is an extra process and network hop; the same request optimizations could run in a client.\n\n## Using Jevaro from Python and JavaScript\n\nTo try it with [uv](https://docs.astral.sh/uv/guides/tools/), set your TypeSafe API key and start the server:\n\n```\nexport TYPESAFE_API_KEY=\"your-api-key\"\nuvx jevaro-server\n```\n\nLeave that running. Save this as `example.py`, then run `uv run --with jevaro example.py` in another terminal:\n\n``` python\nfrom jevaro import Noul, TypeSafeClient\n\nwith TypeSafeClient(base_url=\"http://127.0.0.1:8000\") as client:\n    with client.system_one(\n        states=[\"Please refund the shoes.\", \"Where is my parcel?\"],\n        questions={\"refund\": Noul(instructions=\"Is a refund being requested?\")},\n    ) as reader:\n        print(reader.schema)\n        for batch in reader:\n            if batch.num_rows:\n                print(batch.to_pylist())\n```\n\n`reader` is a PyArrow `RecordBatchReader`. This prints the schema and each state’s refund probability. To collect a PyArrow table instead, replace the loop with `table = reader.read_all()` inside the context manager.\n\nFor JavaScript, run `npm install jevaro`, save this as `example.mjs`, and run `node example.mjs`:\n\n``` js\nconst client = new TypeSafeClient({ baseURL: \"http://127.0.0.1:8000\" });\nconst reader = await client.systemOne({\n  states: [\"Please refund the shoes.\", \"Where is my parcel?\"],\n  questions: { refund: noul(\"Is a refund being requested?\") },\n});\n\ntry {\n  console.log(reader.schema.toString());\n  for await (const batch of reader) {\n    for (const row of batch) console.log(row.refund);\n  }\n} finally {\n  await reader.cancel();\n}\n```\n\nThe JavaScript client returns an Apache Arrow `AsyncRecordBatchStreamReader`. The SDK also works in browsers.\n\n## Throughput and cost\n\nTo measure Jevaro’s throughput, we tested a batch of 10,000 synthetic customer messages, each evaluated with a Choice for department, a Score for urgency, and a Noul for whether a refund was requested. Our fastest run returned all 10,000 rows in **21.5 seconds**, about **464 states per second**. The client received the schema after 34 milliseconds and its first answer row after 290 milliseconds. The elapsed time includes upstream calls, retries, and receiving the complete Arrow result locally; saving it to disk happens afterward. The TypeSafe API’s rate limits are changing often right now, so your throughput may differ. Based on reported token usage and [TypeSafe’s published pricing](https://docs.typesafe.ai/models), the estimated cost was **$0.20 for 30,000 answers**.\n\nJev handled this workload inexpensively and relatively quickly, given the overhead of 10,000 separate API calls. That doesn’t yet tell us how much faster it could be with a bulk API that returned Arrow directly. Jevaro still makes one API call per state, and the TypeSafe API serializes each response as JSON. Jevaro parses those responses and builds Arrow arrays. We’ve optimized that conversion and moved it out of user code, but it still happens.\n\nTo remove that overhead, batching and Arrow output need to happen in the TypeSafe API itself. We’d love to test a native batch endpoint with the TypeSafe team and measure the difference. In the meantime, we’ll keep refining Jevaro: adding Arrow input, improving how it adapts to evolving API rate limits, and experimenting further with output record batch sizes. What would you want from an Arrow interface to Jev? [Open an issue](https://github.com/columnar-tech/jevaro/issues) and tell us.\n\nTo be clear, our wish list for the TypeSafe API is narrow: a bulk endpoint that returns Arrow. Everything else about Jev has us excited, especially the prospect of pipelines that mix probabilistic decisions with deterministic data processing. We’re building some novel capabilities along those lines into [Columnar Gateway](https://columnar.tech/gateway/).\n\n## Next steps\n\n- Try [Jevaro](https://github.com/columnar-tech/jevaro) with your own states and questions.\n- [Sign up for early access to Gateway](https://columnar.tech/gateway/) .", "url": "https://wpnews.pro/news/what-if-jev-spoke-arrow", "canonical_source": "https://columnar.tech/blog/what-if-jev-spoke-arrow/", "published_at": "2026-09-29 16:25:09+00:00", "updated_at": "2026-09-29 16:48:34.716622+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-products", "developer-tools", "structured-data"], "entities": ["TypeSafe AI", "Jev", "Apache Arrow", "Outlines", "Reinforcement Learning for Calibrated Decisions"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/what-if-jev-spoke-arrow", "markdown": "https://wpnews.pro/news/what-if-jev-spoke-arrow.md", "text": "https://wpnews.pro/news/what-if-jev-spoke-arrow.txt", "jsonld": "https://wpnews.pro/news/what-if-jev-spoke-arrow.jsonld"}}