Project: **Baize** β a team-facing AI assistant runtime (Go 1.25+, MIT)
Repo: [https://github.com/rebornace/baize](https://github.com/rebornace/baize)
Before we start, four names in this article map directly to the product UI, so let's align them:
| Name | What it does |
|---|---|
| Main model | Writes replies, calls tools. This is where most of your token spend lives. |
| Decision model | System One (default tev1 ). Only answers short yes/no or pick-which questions: is memory extraction worth it this turn, which backends does this turn need, should a huge tool result be kept, should routing be light or heavy. |
| Fallback model | Configured via decide_profile_id . Only kicks in when the decision model is unavailable.If you've configured a decision model, leave this empty. Leaving it empty doesnot fall through to the main model. |
| Retrieval model | e.g. bge-m3 . Only handles tool retrieval β it does not make any of the judgments above. |
Once you hang Baize next to a business system and point it at an API doc, hundreds of endpoints become hundreds of callable tools. One real backend easily brings in 300+ tools.
With that many tools, every turn carries a pile of high-frequency judgments before the main model ever writes a word: which systems should the main model see, which tools make the candidate list, is memory extraction worth running at all, which model tier to pick.
These judgments share three traits: high frequency (almost every turn), short answers (usually "pick a few" or "yes/no"), and not expensive β if every judgment goes through the main model's generative call, the tokens saved cost more than the judgment itself.
The standard move is to pull high-frequency judgments out of the generative loop and run them on a cheaper capability β the same shape Jev popularized: short answers, enum output, no long reasoning essays. Here's how Baize layers prefiltering, retrieval enhancement, and the decision chain.
This tier needs no decision model. Baize starts with a zero-cost, degradable deterministic path:
It's deterministic, explainable, and works offline. It only pushes down noise though β when the user phrases something vaguely, irrelevant entries can still slip through.
Note: this "narrowing" is a runtime prefilter strategy. It is not the same code as the Rules tier in the decision chain below β Rules today mainly covers "is memory extraction worth running this turn."
v0.4 added the "can you find the right tool" dimension. The prefilter was BM25; dense vectors are now fused via RRF. You can optionally plug in an embedding model (bge-m3 on local Ollama, or any OpenAI-compatible embedding API) to vectorize every tool doc and rank by similarity at query time.
In the fusion, vectors outrank BM25 (in the code, dense weight is roughly 2Γ BM25). It's steadier for mixed Chinese/English, business jargon, and non-standard phrasing. If enhancement fails, it automatically falls back to standard matching β your conversation is unaffected.
This tier answers "can you find the right tools." It does not answer "should we extract memory, should a tool result be kept verbatim, should routing be light or heavy" β that's the decision layer's other track.
v0.5.0 routes "needs judgment" into the decision model, forming today's chain:
Decision model (System One) β Fallback model (optional) β Rules
Tier 1: Decision model (System One). Local Ollama (β₯0.35) auto-pulls tev1 (CPU-friendly), or you can point it at any cloud / self-hosted service compatible with POST /v1/systemone (not a plain Chat Completions endpoint). Short, high-frequency judgments are answered locally or on a dedicated decision endpoint β they don't burn the main model's generative budget.
Tier 2: Fallback model (optional). Only when System One is unavailable, a cheap preconfigured model profile takes over. If you've already configured the decision model, leave this field empty β leaving it empty does not route back to the main model; it falls through to the next tier.
Tier 3: Rules. Deterministic rules. Today this mainly covers memory extraction; other judgment points opt out of this tier and continue on the original path: we'd rather follow a hard-coded fail direction than pretend to be smarter than we are.
The whole chain is fail-open: if a tier times out, isn't configured, or opts out, it doesn't drag the conversation down β tools schemas are still sent, memory extraction still runs. We'd rather spend one extra call than silently drop a capability.
About scores. The System One protocol may return an internal ranking score; the client only uses it to pick an enum result (e.g. β₯0.5 β yes). It does not expose scores to product logic, and it never writes "if p > 0.9, auto-execute." The numbers a typical local decision model emits are uncalibrated ranking scores, not business thresholds. Truly calibrated probabilities ("90% means ~90%") require a dedicated training path; the open-source default path doesn't pretend to have it.
Why now? The decision model wasn't invented in v0.5. v0.4 already pulled local Ollama installation, bge-m3 pulls, model paths, and uninstall into the "Matching & Decision" settings page. v0.5's delta is mostly pulling one more model (tev1) on the same Ollama, behind the same install UX, unified under /v1/systemone. Once the infrastructure is in place, the marginal cost of adding a decision model is low.
Why CPU-friendly? High-frequency judgments need to be short and resident. tev1 is CPU-friendly per the product docs β it runs on an ordinary office machine, so you don't round-trip a cloud generative call just to "think for a second."
Why two layers of switches? If you've configured a decision model but haven't flipped the master toggle under "Runtime β Smart acceleration," you do not save main-model tokens. After the master is on, you still turn on the individual features you want: tool candidate narrowing, tool-result pruning, memory pre-judgment, long-doc selection. The system needs to know explicitly which capabilities are on β it doesn't assume "model installed = everything on."
Retrieval and decision can share one Ollama. bge-m3 for matching and tev1 for decision can both live on the same Ollama instance. The "Matching & Decision" settings page keeps both links together: install progress, model paths, custom directory, cleanup β all in one place.
Installing Ollama. Both System One and enhanced matching assume you have Ollama running. The settings page walks you through install, model pull, and cleanup β get Ollama first, none of the decision chain matters if you can't run it.
Don't mix "we swapped the decision backend" with "the main model's prompt got shorter" on one chart.
Benchmark: 37 real read-only business requests across 3 backends (390 tools total), main model DeepSeek-Flash. Relative to prefilter width 32, the default width 16 cuts turn-0 prompt tokens by ~34%; over 5 rounds (185 runs), run success is 99.5%. Compared to sending all 390 tools in one turn (~85k prompt tokens at turn-0 vs ~1.4kβ3.8k), it's an order-of-magnitude drop, not just 34%.
These numbers reflect tool narrowing + prefiltering on the main model's prompt. Corpus and scripts live in scripts/tool-routing-eval and are reproducible. It is not "System One alone saves 34%" β when prefilter width is fixed and only the decision backend swaps to tev1, the main model's prompt width is roughly the same.
A decision-model call is usually far cheaper than a main-model call (self-hosted tev1 β electricity). What you actually want to see is: which processes no longer go (or go less) through generative calls.
| Process | Without smart acceleration / decision model | With System One on and the matching toggle enabled |
|---|---|---|
| Run memory extraction? | Often another generative call to extract | Decide first; if "not worth it," the whole extraction is cancelled |
| Which backends this turn? | Full schema (or another model asked) | Decision + prefilter; main model only sees the narrowed list |
| Keep a huge tool result? | Often the whole block goes into context / summary | Decide first; if useless, don't feed it to the main model |
| Light or heavy routing? | Heuristic, or another model arbitrates | Optionally decided by the decision model |
The big savings are: fewer schemas and less context fed to the main model, plus generative calls cancelled outright. Swapping the decision backend for a cheaper implementation is the second ledger.
Same 37 requests, width 16, same machine, one run each (this is a decision-backend comparison β not another pass over the width table):
| Decision backend | Run success | Expected tools hit | Wall clock |
|---|---|---|---|
| System One (local `tev1` ) | 37/37 | 31/37 | **~376 s** (~10.2 s/request) |
| Fallback model ( `deepseek-flash` ) | 37/37 | 29/37 | ~556 s (~15.0 s/request) |
Event sources are systemone and remote; main-model prompt width is the same. The fallback model's wall clock is ~1.48Γ β wiring in a decision model isn't "hiring another model to chat along," it's pulling high-frequency judgments out of generative calls. Scripts and notes are in scripts/tool-routing-eval.
Requirements: Go 1.25+, or grab a prebuilt binary from GitHub Releases (use v0.5.0). Out of the box it's still zero-component standard matching. For the full decision chain:
tev1, or point at /v1/systemone).
The article tracks the repo; the repo docs are the source of truth.
From deterministic prefiltering, to optional vector retrieval, to a decision model wired into a chain β Baize moves high-frequency judgments from expensive places to cheaper ones, and when it can't or isn't sure, it fails open onto the original path. Judgment can plug into a dedicated decision endpoint; whether it's accurate and how well-calibrated it is, that's another ledger. Try it and tell me what breaks.