# I don't watch F1, so I built a quiz where the AI isn't allowed to grade me

> Source: <https://dev.to/hempun10/i-dont-watch-f1-so-i-built-a-quiz-where-the-ai-isnt-allowed-to-grade-me-b0b>
> Published: 2026-10-02 07:48:29+00:00

*This is a submission for the [Sanity Challenge, Path One: Ship an Agent That Queries Real Content](https://dev.to/challenges/sanity-2026-09-16)*

I'm not an F1 person, but I've wanted to get into it for a while. Pole, DNF, sprint, undercut: I knew the words existed and not much more. So for this challenge I built the thing I actually needed, a quiz that teaches you Formula 1 one race at a time.

It's called Pit Wall. You pick a 2026 Grand Prix and a tyre (Soft for one fact from the race, Medium for comparing drivers, Hard for the season so far), the five start lights go out, and you race five laps of questions. After every answer the timing tower slides from grid order into the finishing order, a card explains the F1 term the question used, and you can radio your race engineer to ask why.

The easy way to learn would be to ask a chatbot. The problem is that a chatbot answers 2026 questions quickly and confidently, and often wrongly (I measured it, below). So the rule for Pit Wall is simple:

**The AI writes the questions. It is never allowed to decide if you got them right.**

An agent writes each question together with the GROQ query that answers it. My server runs that query against the race data in Sanity and keeps whatever comes back. The model's opinion of the answer is thrown away.

To check this mattered, I took 30 questions from the app and gave them, with the same four options, to the same model with no access to the data:

|  | Correct | 
|---|---|
| Pit Wall (answer from a fresh GROQ query) | 30 / 30 | 
| Same model, no data | 16 / 30 | 
| Same model, Hard questions only (season so far) | 3 / 10 | 

Asked who had the most podiums after Monaco, it said Max Verstappen. The data says Kimi Antonelli. It made the same kind of mistake on almost every season-level question. Random guessing would get 7 or 8 of the 30. So 16 beats guessing, but it's still wrong almost half the time, and 7 times out of 10 on the season questions. That's the number that made the extra plumbing feel worth it.

**Play it:** [https://pit-wall-olive.vercel.app](https://pit-wall-olive.vercel.app) (no sign-up, pick a three-letter driver code and race)

The README covers the setup from an empty Sanity project, plus credits for the data, fonts and sounds.

The race results come from [F1DB](https://github.com/f1db/f1db) (CC BY 4.0). A small Python script turns their SQLite release into Sanity documents: races, results, drivers, teams, circuits and championship standings from 2014 to round 15 of 2026.

One result document is one driver in one race, with typed fields for the things a newcomer asks about: `gridPosition`, `position`, `positionsGained`, `pitStops`, `polePosition`, `fastestLap`, `points`. Each result references its `race`, `driver` and `team`, so a question like "which podium finisher gained the most places at Monza" is one GROQ query instead of a pile of text the model has to read and add up.

A couple of fields have notes because the data surprised me:

`points` is Grand Prix points only. Sprint points live in the standings documents, so the agent is told never to call a sum of `polePosition == true`, not `gridPosition == 1`. Those differ whenever someone takes a grid penalty.` pitStops: 0`. In a dry race that's not possible (you have to run two tyre compounds), so it's missing data, and the agent is told not to build questions on it.
Pit Wall uses two Sanity Context MCP endpoints:

`pit-wall-data` points at the dataset in GROQ mode, for facts and numbers.`pit-wall-rules` points at a Knowledge Base, for what the words mean.
Sanity's docs warn that if one endpoint has both a dataset and a Knowledge Base, the dataset wins and the Knowledge Base is ignored without an error. So there are two, and the agent gets both tool sets with a prefix on each, because both endpoints have a tool called `initial_context`.

One thing that cost me time: the data endpoint kept returning an error until I ran `sanity deploy`, because it only works for a dataset with a deployed Studio.

The rules come from nine Wikipedia articles (points systems, race weekend format, tyres, safety car, DRS, the 2026 season, a glossary), imported with `sanity context imports create` and built into 15 entries.

I also imported one stale source on purpose: a June 2023 version of the race weekend article, which still says the fastest lap earns a bonus point. That rule was scrapped from 2025.

The Knowledge Base build caught it. It flagged four conflicts, and one was exactly this: one entry said the point was abolished, the glossary said it still exists. What I liked is that the race data could settle it on its own. I checked every race winner who also set the fastest lap:

| Season | Points for a win plus fastest lap | 
|---|---|
| 2018 | 25 | 
| 2019 | 26 | 
| 2024 | 26 | 
| 2025 | 25 | 
| 2026 | 25 | 

So the bonus existed from 2019 to 2024 and is gone now. I resolved the conflict to "abolished from 2025", and that became a standing instruction for the Knowledge Base.

The other three flags were a mix. Q3 being 12 or 13 minutes long came from the same old article (it's 13 from 2026). Tyre compounds C1–C5 versus C1–C6 had no source on one side. And one claimed Antonelli couldn't have been the last driver to use DRS because he raced in 2026, which misread a date range, so I dismissed it.

`groq_query` on the data endpoint to find something worth asking and writes the query that answers it.`submit_question` tool.
Every answer screen has a "How we checked this" drawer that shows the exact query that graded you.

The first version wrote every question live. It took 20 to 50 seconds and roughly 40,000 to 55,000 input tokens each, and a single page load could hit my OpenAI limit of 200,000 tokens a minute.

What helped, in order:

`quizQuestion` document.
Now most plays come straight from Sanity in about a second and cost nothing. The agent only writes a new question when you've seen everything for that round and tyre, and it stops at six per round and tyre, which caps what the app can ever spend. There are 222 verified questions in the bank right now, and you can browse them in Studio with the query that proves each one.

After you answer, you can radio your race engineer. This one is a live agent with both endpoints. It answers in a few sentences, and under every reply there's a "What I checked" list with the queries it ran and the rules entries it read.

It's honest about its limits. When I asked why a driver made only one pit stop, it confirmed the number and then said the data doesn't record the reason. Ask it for a poem and it says it can only talk racing on this channel. It costs under a cent a message, and there's a limit of two per lap.

A finished race scores like F1's top five: 25, 18, 15, 12, 10 for five to one correct. Your best result for each race and tyre counts, and ties go to whoever got there first, like identical lap times in qualifying. Finished races are saved as `quizRun` documents and the standings are worked out from them.

There are no accounts, just a three-letter code like a driver. To keep the scores honest, your score travels inside an encrypted token from answer to answer and the server reads it from there. I tried replaying a question, forging the token and finishing the same race twice, and each was rejected or saved once.

The question bank repeats itself. "Who has the most podiums after round X" shows up a lot on Hard, partly because Antonelli has won 8 of the 15 races so far, so the answer barely changes. Next time I'd push the agent toward a wider mix of question types.

No accounts means duplicates. I'm on the standings twice as HEM because I played from two browsers. Optional sign-in would fix it; I left it out so nobody has to sign up before the lights go out.

The Knowledge Base is small. Nine articles cover what a newcomer asks, not race strategy.

I built Pit Wall with Pi: the data import, the agent, the app, the design (a warm-paper "FIA timing sheet" look with a pixel font), the sounds and the deploy. The F1 car you hear at lights out is a real recording by Geoff-Bremner-Audio (CC BY 4.0). The fonts are Departure Mono and Instrument Serif, and the flags are from country-flag-icons.

`hptt7wjq`
`production` (public)`race`, `raceResult`, `driver`, `team`, `circuit`, `driverStanding`, `teamStanding`, `quizRun`
I built this because I wanted a way into F1 that doesn't make things up. Two questions for the comments:
