Updated
Evaluating how models perform across a range of knowledge work tasks, using live anonymized traffic and judged by models from Anthropic, OpenAI, and Google. An ongoing experiment.
This is not a traditional benchmark. Benchmarks are stable and curated sets of difficult, fully specified tasks, which are designed to discover new capabilities as models advance. Notion’s Knowledge Board asks a different question: How do models handle the tasks users actually delegate to them?
Notion users rely on agents for all kinds of knowledge work: meeting follow ups and action items, inbox and calendar cleanup, support-ticket follow-ups, sales-pipeline updates, lead research, knowledge-base Q&A, recurring project reports, and more. Our Knowledge Board is a live, ongoing measurement of how well models in Notion do that work.
We developed an ensemble judge in collaboration with researchers at Anthropic and OpenAI. The chart below traces the path end to end: a small, random slice of live Notion traffic is split evenly across models, each outcome is scored by an ensemble of independent judges, and the surviving traces become the metrics shown throughout.
Many of the differences we measure are smaller than their confidence intervals, so we’ve deliberately avoided creating a leaderboard. The data on this site is presented to make that visible. Let us know how we can improve: [email protected].
The charts below offer two ways to read the same data. The first compares models on one metric at a time. (Note that overlapping confidence intervals mean models tie; our Knowledge Board offers a live snapshot of model performance, not a leaderboard.) The second chart plots quality against cost, so you can spot the models that deliver the most resolved work for the least money or time.
Metric estimates #
Each model on one metric, with confidence intervals.
Model performance in Notion, based on percentage of knowledge work tasks resolved across CX, project management, and more.
| Model | Estimate | 95% CI lower | 95% CI upper |
|---|---|---|---|
| Opus 5 | 98.4% | 98.2% | 98.6% |
| Kimi K3 | 98.2% | 98% | 98.4% |
| Opus 4.8 | 98% | 97.7% | 98.3% |
| GPT-5.6-Sol | 97.4% | 97.1% | 97.6% |
| Grok 4.5 | 97.2% | 96.9% | 97.5% |
| GPT-5.5 | 97.2% | 96.8% | 97.6% |
| Opus 4.7 | 96.8% | 96.4% | 97.3% |
| Sonnet 5 | 96.7% | 96.3% | 97.1% |
| GPT-5.4 | 96.6% | 96.2% | 97% |
| GLM-5.2 | 96.5% | 96.2% | 96.8% |
| GPT-5.6-Terra | 95.5% | 95.2% | 95.9% |
| GPT-5.6-Luna | 95.3% | 94.8% | 95.8% |
| Sonnet 4.6 | 94.6% | 94.1% | 95.2% |
| Kimi K2.7 Code | 94.1% | 93.5% | 94.7% |
| DeepSeek V4 Pro | 93.7% | 93.1% | 94.2% |
Resolution rate, 95% CI
Trade-off frontier #
Quality vs. cost in terms of dollars or seconds—up and left is better.
Cost per task in USD for models to resolve knowledge work tasks.
| Model | Resolution rate estimate | 95% CI lower | 95% CI upper | $/task estimate | 95% CI lower | 95% CI upper | Probability of being on the frontier |
|---|---|---|---|---|---|---|---|
| Opus 5 | 98.4% | 98.2% | 98.6% | $0.87 | $0.86 | $0.89 | 85% |
| Kimi K3 | 98.2% | 98% | 98.4% | $0.45 | $0.44 | $0.46 | 100% |
| Opus 4.8 | 98% | 97.7% | 98.3% | $0.71 | $0.68 | $0.74 | 13% |
| GPT-5.6-Sol | 97.4% | 97.1% | 97.6% | $0.64 | $0.61 | $0.67 | 0% |
| Grok 4.5 | 97.2% | 96.9% | 97.5% | $0.47 | $0.45 | $0.50 | 6% |
| GPT-5.5 | 97.2% | 96.8% | 97.6% | $0.62 | $0.58 | $0.65 | 0% |
| Opus 4.7 | 96.8% | 96.4% | 97.3% | $0.70 | $0.66 | $0.74 | 0% |
| Sonnet 5 | 96.7% | 96.3% | 97.1% | $0.34 | $0.32 | $0.35 | 76% |
| GPT-5.4 | 96.6% | 96.2% | 97% | $0.38 | $0.36 | $0.40 | 36% |
| GLM-5.2 | 96.5% | 96.2% | 96.8% | $0.17 | $0.16 | $0.19 | 100% |
| GPT-5.6-Luna | 95.3% | 94.8% | 95.8% | $0.02 | $0.02 | $0.03 | 100% |
| GPT-5.6-Terra | 95.5% | 95.2% | 95.9% | $0.30 | $0.29 | $0.31 | 0% |
| Sonnet 4.6 | 94.6% | 94.1% | 95.2% | $0.39 | $0.37 | $0.41 | 0% |
| Kimi K2.7 Code | 94.1% | 93.5% | 94.7% | $0.17 | $0.16 | $0.19 | 0% |
| DeepSeek V4 Pro | 93.7% | 93.1% | 94.2% | $0.28 | $0.27 | $0.30 | 0% |
$/task
Methodology #
Context
Our goal with this work is to give our customers a helpful resource for deciding which model best fits their work, along with a tool for understanding the trade-offs between cost, average resolution time, and performance.
Many rigorous benchmarks already exist, but none capture the full range of industries that Notion customers use Notion AI for, nor do they cleanly reflect the distribution of tasks that people delegate to the agent inside Notion.
So we built the Notion Knowledge Board. It evaluates models in Notion’s production environment, using real user requests and prompt shapes, to cover the broad knowledge-work use cases that happen in Notion. We plan to maintain the Knowledge Board and update it as new models are released and as users discover new capabilities and task distributions shift.
Random assignment
First, an experiment gate evenly and randomly splits eligible live agent traffic across the models we’re evaluating. We did this because we believed using data from users that self-select a model would introduce bias in types of requests and user distribution, and we wanted to avoid any such biases.
Models in the experiment
We used the following models in the experiment:
- DeepSeek V4 Pro (high)
- GLM-5.2 (high)
- GPT-5.4 (high)
- GPT-5.5 (high)
- GPT-5.6-Luna (high)
- GPT-5.6-Sol (high)
- GPT-5.6-Terra (high)
- Grok 4.5 (high)
- Kimi K2.7 Code (thinking)
- Kimi K3 (high)
- Opus 4.7 (high)
- Opus 4.8 (high)
- Opus 5 (high)
- Sonnet 4.6 (high)
- Sonnet 5 (high)
Only the models that passed our agent baseline private evals with a memory score of 0.8 or higher were considered as candidates. Models that did not pass our Zero Data Retention requirements, including Fable, were excluded. Models were also excluded if they were found to have high error rates, or if we discovered bugs while scoring them.
All models were tested on “high” effort level, or the closest equivalent. (Note that “high” reasoning effort does not represent a consistent reasoning envelope between vendors or models.)
Opt-outs and exclusions
No customer data was retained during this experiment. The resolution score and product metrics were calculated online, and only the 0/1 flag was saved. Users could also opt out of the experiment at any time, and change the model they use. If so, their requests were excluded from the data. We also excluded requests if there was any kind of harness error, or if the judge decided that the request is not fair to be evaluated.
What we collect
We developed an ensemble judge in collaboration with researchers at Anthropic and OpenAI. The ensemble judge assessed the back-and-forth of messages between the user and the agent as well as the tools the agent called and any actions it took; decided if the user request was successfully met; and then decided if the failure was a fair case. We used internal Notion data to align the judge and test its performance with the assistance of human annotators.
We then collected the resolution rate as a true/false signal from the experiment cohort. We did not collect any additional PII or user data alongside the final grade.
In addition to resolution rate, we used our anonymized product metrics to form a holistic view on model performance:
- Average resolution time (s): total seconds taken from when a user submits their request until the agent finishes responding.
- Resolution rate (%): percentage of requests completed or handled appropriately, including necessary follow-up questions and valid refusals.
- Average cost ($/task): average dollars for a model to resolve one task.
- Undo rate: % of agent responses undone by users.
Statistical treatment
Because models were randomized per user, and the randomization is not enforced at request level, requests from the same user are correlated. To account for this correlation, we obtained the variance of a ratio metric R = Ȳ/X̄ using the delta method:
Var(R) ≈ (1/n) · ( σ2Y − 2R·σXY + R2·σ2X ) / μ2X We collected enough data to be able to detect a 1% difference in resolution rate per model, compared with our baseline. (We are not going to share comparisons or lift numbers, as we don’t expect to declare a winner.) The ranking of the models is based on the point estimates, and viewers should use the noted confidence intervals to decide if the difference between two models is statistically significant.
Implementing the LLM judge
When a user request resolves, we present the transcript to three different models: Anthropic’s Claude Opus 4.8, Google’s Gemini 3.5 Flash, and OpenAI’s GPT-5.5. Each model is given the prompt below, verbatim.
Each judge returns results mapped to either one of true | false | null
based on whether the user’s request was resolved. If one judge decided that the trace is not evaluable, that trace was excluded from results. Traces are given a final score of 1 or 0 based on the majority vote of the judges.
| Final Resolved Tag | Judge values |
|---|---|
| True | COMPLETED, APPROPRIATE_CONTINUATION, APPROPRIATE_REFUSAL |
| False | FAILED |
| Null | NOT_EVALUABLE, NOT_APPLICABLE |
This judge prompt was iterated on with Anthropic and OpenAI, both of whom reviewed a draft of the judge prompt and provided feedback on its design.
Our research team also conducted a round of internal expert human labeling on 100 transcripts, producing a κ score of 0.595 and overall agreement of 92.9%. The majority of disagreements were human-labeled failures with ensemble-graded passes, suggesting that the judge prompt is slightly more lenient.