An agent picks up a task, works on it, marks it completed. You look at it the next morning.
How long did it take? How much of that was the agent working, and how much was it waiting for you to click Allow?
For a long time AgentRQ (an open-source task board and MCP server for Claude Code and other agents) couldn't tell you. A task had one status field, and every change overwrote it. ongoing → blocked → ongoing → completed, and each write erased the one before. The end state was right. Everything in between was gone, and you can't get a duration out of a snapshot.
We fixed that in agentrq/agentrq#756. This post isn't a code walkthrough. It covers the design decisions behind it, because I think they apply to anyone building agent tooling: why the history is immutable, why the states are a fixed enum, why there is exactly one writer, and why the most useful state in the timeline is one the task can never actually be in, and how it all rolls up into a p50 dashboard.
TL;DR
Because the interesting question about agent work is almost never "is it done?" It's "where did the time go?"
With a human ticket, you can ask the assignee. With an agent, the only record is what the system kept. If the system kept one mutable field, you have a photograph, not a film.
It's the event-sourcing lesson in miniature: the current state is a projection; the transitions are the data. Keep the transitions and you can always work out the state. Keep only the state and the transitions are lost for good.
We didn't go full event sourcing, though. The task's status stays the source of truth for "what is it now", because a lot of the system reads it: the work queue agents pull from, filters, counts, and the MCP tools agents call. We didn't want to touch any of that. The history sits beside it.
Every status change appends one record: the state the task came from, the state it went to, when, and which agent (and model) made the change. That's it. A record is never edited after it's written, and nothing is ever inserted into the past.
not started 09:00 ─┐ 6m
ongoing 09:06 ─┤ 24m
needs input 09:30 ─┤ 8m
ongoing 09:38 ─┤ 31m
blocked 10:09 ─┤ 38m
ongoing 10:47 ─┤ 19m
completed 11:06 ─┘
A state lasts from its record to the next one. The state the task is in right now lasts until "now", so it keeps counting.
Two honest caveats about "immutable": if you move a task to another workspace, its history moves with it, and if you delete a task, its history goes too. What never changes is the content of a record: which states, and when.
Three reasons, in order of how much they bit us.
1. The numbers have to be trustworthy, or they're useless. The whole point of the timeline is to answer questions like "is this class of work worth handing to an agent?" and "am I the bottleneck?". If records could be edited, every number derived from them would be an opinion. Append-only means the only way to change the story is to add to it.
2. Agents write status too, and they're not always right. An agent might mark a task completed, then get reopened by a person, then complete it again. With an overwritten field, the final completed is all you see. With an immutable history, the false finish is on the record, with its time. That's exactly the kind of thing you want to spot when you're tuning a prompt or a workflow.
3. It makes concurrent writers boring. An agent and a person can move the same task at the same moment: the agent marks it completed while you click Reject. With an append-only log, there's nothing to merge. Both changes are recorded in the order they happened, and each record says where it came from. (We do lock the task while recording, so the second change sees the first one's result as its starting point.)
There's a fourth, quieter benefit: the conversation got cleaner. Every status change used to drop a "Status updated to …" line into the task thread, which buried the messages that mattered. With the history in its own place, those lines are gone.
This looks like a small decision. In an immutable log it's a big one.
A record written today has to mean the same thing forever. Free-form strings drift. Someone writes complete instead of completed, a state gets renamed, a new client sends in_progress. In a mutable column you can fix that with a migration. In an append-only history, "fixing" old records is exactly what you promised never to do. So the set of states is closed: not started, ongoing, blocked, needs input, completed, rejected, and scheduled (cron). Nothing outside that set can be recorded as itself.
The enum only ever grows at the end. Each state is stored as a small number, and new states are appended after the existing ones, never inserted in between or reordered. Reordering would silently change the meaning of every record already written, so it's the one rule that can't be broken. "Needs input" was added exactly this way, as the last value.
There's an explicit "none" state. When a task is created, the first record goes from "none" to its first status. That gives the first segment a start time, and it gives any unrecognised status somewhere safe to land instead of failing the write.
Numbers inside, names outside. Storage uses the compact number, because this history grows with every status change on every task and the workspace stats pages aggregate over it. Everything outside the backend, including the API, the UI and the agents, sees the readable names. The mapping lives in one place.
Because a history that some code paths forget to write is worse than no history: it looks complete and isn't.
Status changes come from everywhere: the MCP tools agents call, the web app, the CLI, scheduled tasks, events that create tasks in other workspaces. Instead of asking each of those to remember to record a transition, the record is written in one place, the storage layer's task save, and in the same transaction as the status change itself.
That gives two guarantees:
completed with no trace of how it got there.
Everything above the storage layer sets the status and saves, as it always did. It doesn't know the history exists.
This is the part I'd steal if I were building something similar.
When an agent asks the person something through AgentRQ (a permission request before running a command, or a form question), the task is waiting on the human. That's a different kind of wait from blocked, where the agent hit a wall and said so. They call for opposite fixes, so they deserve separate numbers.
The obvious implementation is a new status. We didn't do that. Agents read and write the status, the work queue hands out tasks based on it, and filters and counters group by it. A new value would change behaviour in all of those places: an agent looking up its ongoing tasks could suddenly miss the one it is in the middle of.
So "needs input" is derived, and it only exists in the history:
ongoing the whole time.
A few rules made it correct:
Because the status never changes, nothing that reads it behaves any differently. Only the timeline and the totals know.
Under each task's description there's now a task timeline: one dot per state, with the time spent in each on the line between two dots, and the totals above.
| Total | Value | Meaning |
|---|---|---|
| Worked | 1h 14m | All the time the agent was ongoing |
| Blocked | 38m | All the time the agent said it couldn't go on |
| Needs input | 8m | All the time a question to the person was open |
| Start to close | 2h | From the firstongoing to completed or rejected |
Start to close begins at the first ongoing, not at creation, on purpose. A task that sat in the queue overnight shouldn't look like an agent that took twelve hours.
An open task keeps ticking. Here the agent is waiting on a permission prompt, so the last dot is yellow and counting, and there's no start to close yet:
A small consistency detail: the page works out the totals from the same segments it draws, instead of showing the server's numbers next to a bar it animates itself. Otherwise the live counter and the bar would drift apart every second. The server's totals are still in the API for scripts, agents and the stats pages.
On a phone it collapses to one line, the dots and the total:
One task's timeline tells you what happened to that task. It doesn't tell you whether that was normal. Is two hours usual for this workspace? Did this week go faster than last? Do agents sit blocked longer in one workspace than another?
So the history feeds a second view: a Task Latency panel on each workspace's stats page and on the account stats page. It charts the same four totals, in minutes, over every task closed in the range you pick, with a toggle between P50 (the typical task), Min (the quickest) and Max (the slowest).
Read the cards like a story. This workspace closed 42 tasks in the week. At p50, a task took 108 minutes start to close, of which 66 were work, 19 were blocked and 7 were waiting on a person. (Each line is its own median, so they don't add up exactly, and they aren't meant to.) Hover over Sep 24 and the typical task took 150 minutes, well above the week's figure. That's the day to go and look at.
The account page shows the same panel across every workspace:
Across the account the typical task took 146 minutes and was blocked for 26 of them. The workspace above beats both, so something else is pulling the account figure up, and you know where to look next.
The design decisions behind the numbers:
Why p50 and not the average? Agent task durations have a long tail. One task that sat blocked over a weekend turns a 30-minute average into a 3-hour one, and then the average describes no task you actually ran. The median is the task you'd typically get. Min and max sit next to it on purpose: max finds the one that dragged, and min shows how fast the work goes when nothing gets in the way.
Why an estimated p50? Stats are rolled up per hour, per day and per month, so a three-month chart doesn't re-read every task. But medians don't add up. You can't combine two hours' medians into a day's median. So each rollup keeps a small histogram of durations with fixed, log-spaced buckets: under a minute, 1–2 minutes, 2–5 minutes, and so on through hours and days, up to over 30 days. Histograms merge by adding counts. The p50 is read back out of the merged histogram, interpolated inside the bucket where the median falls, and clamped between the exact min and max, which merge perfectly. It's an estimate, and it's exact when a bucket holds a single task.
The histogram buckets are immutable too. Same lesson as the state enum: counts are stored by position, so the bucket edges can never be changed or reordered. A finer histogram would be a new field, not an edit.
One arithmetic, two views. The task page and the dashboard compute durations with the same function over the same history, so a task's timeline and its contribution to the chart can never disagree.
Every close counts once. When a task closes, its totals are recorded in the same transaction that closes it. If it's reopened and closed again, its totals are recalculated over its whole history, but a period that was already rolled up isn't rewritten. Each close counts once, in the period it happened in.
Zero is data; nothing is a gap. A task that was never blocked counts as 0 minutes blocked, so the blocked line shows how much blocking a typical task really sees, not just the unlucky ones. But a day with no closed tasks is left as a gap in the line, not drawn as zero, so a quiet day never looks like a fast one.
The resolution follows the range. A day or two plots by the hour, up to about three months by the day, and anything longer by the month. All in UTC.
Agents can read it too. The same figures are available to an agent, so you can ask one to watch for a workspace whose blocked time is creeping up, or to report how long its own work took this week.
It shipped as a follow-up, in agentrq/agentrq#758. The full tour is in Task Latency: How Long Your Agents' Tasks Take.
"How long did that take?" finally has an answer. Start to close per task is the number you want when you decide whether a class of work is worth handing to an agent at all.
Agent time vs. your time. A two-hour task might be two hours of work, or twenty minutes of work plus an hour and forty minutes waiting on a permission prompt. The first means the task is big. The second means the human is the bottleneck, and the fix is a broader permission rule, a better prompt, or handing it off at a different time of day. Worked and needs input are separate numbers, so you know which one you have.
Blocked stops being invisible. Before, you saw a blocked task only if you looked while it was blocked. Now every blocked stretch stays on the record with its length, and the dashboard shows which workspace keeps getting stuck.
Improvements become measurable. Change a prompt, a skill or a workflow, and watch whether the typical start to close drops the next week. One fast task proves nothing. A p50 over a week of tasks does.
Why not full event sourcing?
Too much for this. The status is read in many places, including by agents over MCP. An append-only history beside the current status gets the durations without migrating every reader.
Why not work out durations from the message thread?
Messages and status changes are different streams. An agent can go blocked without writing anything, and a quiet agent can be working for 30 minutes. A dedicated history records exactly what changed and when.
Does "needs input" change a task's status?
No. It exists only in the history. Queues, counts, filters and agents see the same status as before.
Why does the dashboard show p50 instead of the average?
Agent task durations have a long tail. One task blocked over a weekend drags the average far from anything typical. The median doesn't move, and max is right next to it for finding the outlier.
Which states can appear?
Not started, ongoing, blocked, needs input, completed, rejected, and scheduled for recurring or one-time scheduled tasks.
The change is in agentrq/agentrq#756, and the Task Timeline feature page has the short tour. Open any task that has changed status since the update and look under the description.
Question for you: if you run agents against a task queue, how do you tell "the agent was slow" from "I was slow"? Do you track human wait time separately, or does it all go into one latency number?
Originally published on the AgentRQ blog.