cd /news/ai-agents/a-practical-guide-to-making-yourself… · home › topics › ai-agents › article
[ARTICLE · art-142795] src=ana15.substack.com ↗ pub= topic=ai-agents verified=true sentiment=· neutral

A Practical Guide to Making Yourself Obsolete Through AI

A developer published a practical guide and open-source GitHub repository (WhiteCollarAgent) showing how to turn white-collar knowledge-work tasks into agent environments, expose tools via MCP servers, sandbox the agent with Docker, and evaluate task correctness, drawing on Mercor's Archipelago implementation, its APEX knowledge-work datasets, and the SkyRL reinforcement-learning framework. The writeup outlines a path toward wrapping these environments into an RL pipeline so agents can independently improve at specified tasks.

by read24 min views1 publishedSep 30, 2026
A Practical Guide to Making Yourself Obsolete Through AI
Image: source

Everyone tells us that AI is getting more powerful. Noam Brown reported that recursive self-improvement may be closer than we thought; OpenAI announced that an AI had solved the Navier–Stokes Millennium Prize Problem and Dario Amodei said we must pace the frontier as progress was moving too fast.

But then OpenAI published an article saying that “fully autonomous RSI is not happening today”, and Anthropic released another article stating that Claude “leads” just 26% of Anthropic’s AI R&D work.

So how close are we, really, to having at least some of the tasks we do at our white-collar jobs be reliably automated by AI?

As noted by Arvind Narayanan, one way AI may impact knowledge work is that “if anything is understood well enough to be specifiable as a task, it can be handed off to AI”. In this future, humans would then remain responsible both for defining those tasks and for doing the work that cannot easily be specified, but potentially freed from the repetitive, less creative work.

Where to start, if you want to be ahead of the crowd and automate some of your tasks, prior to your boss having figured out how to automate away your whole job? In this blog, I’ll show you in detail,

  • how to turn real work tasks into environments an AI agent can navigate,
  • how you can give the AI access to tools similar to those you as a white-collar worker would have using MCP servers,
  • how you can wrap this in an environment the AI can interact with through Docker,
  • how you can judge whether the AI performed the task correctly,
  • and a small comment on wrapping this all into a reinforcement learning pipeline, to have the AI independently learn to get better at the tasks (stay tuned for Part 2, where I’ll discuss the RL in more detail).

This blog is based on:

  • A blogpost by Mercor (one of the world’s largest data providers) on how they trained AI agents to perform frontier knowledge work: [click here]
  • Mercor’s Archipelago agent implementation: [click here]
  • Mercor’s frontier knowledge work tasks on HuggingFace: [click here] (there’s also a newer release which directly allows to run the worlds with Harbor,here )
  • for RL, I’ll be using SkyRL here .

If you want to follow along with everything I do in this blogpost, clone my GitHub repo: https://github.com/abrvkh/WhiteCollarAgent.

Your work world and tasks

Imagine you work as a management consultant, and currently you’re staffed on a project for CompliSure, a company that is considering expanding into Latin America. After getting off a call with the company yesterday, they told you that the Latin American customer count is expected to continue growing at the same rate as last year. Their CEO told you she’s confident that if they decide to enter the region, they can capture a quarter, or even a half, of the market currently held by the two largest players in the area. You left the office yesterday with a half-finished spreadsheet in which you were working on calculating a forecast for 2030. On your way to work, you see a Slack message from the CEO: ‘ping me the numbers when you have them’.

Your task: reply with the numbers.

Your world: in essence your computer, with a bunch of different files, Excel, a mailbox, a calendar, Slack. Inside a folder called Competitive analysis you had some notes on the Geographic_locations_competitors_2016-25B expanded.xlsx and Financial_dataset_2016-25B_expanded.xlsx and inside the folder Forecast model is the pretty-much-done CompliSure_5yr_Forecast.xlsx you need to double-check for the numbers.

Your tools: mostly excel, which you manipulate through its visual interface, and your knowledge of the maths needed to do the calculations.

If we want to teach the AI to get good at this, we need to craft a world with access to the same information (files, excel and so on), give it access to tools it can use to manipulate this world (the AI won’t work through the visual interface; it’ll directly manipulate the file contents), and specify the exact task.

For example, if I were to go about automating my quant researcher work routines, I’d need to create a world with a lot of different datasets, including limit order books from a bunch of instruments, lots of code for properly cleaning and preparing this data, some code for baseline predictors, some code for running backtests on these predictors and a pipeline for doing deep learning on this data. The tasks would be around finding new predictive strategies, that perform good in the backtests and achieve a decent Sharpe ratio.

Experts are needed in order to craft the above tasks. And that is the reason that Mercor is valued at $10B, Scale AI was acquired for $14B, Handshake is at $2B, and Surge AI is around $15B. There is a lot of human effort that goes into creating the relevant worlds, tasks, and golden solutions, and the frontier labs are paying large amounts of money to get access to this data. The market for automating white-collar work is enormous, but human data is still valuable.

Setting up the AI’s world

Having determined the structure of our tasks, we need to replicate this into a world the AI can access. The AI’s world is, similar to ours, an initial state of the workplace. For the management consulting task, it looks roughly like:

world/
├── filesystem/
│   ├── Forecast model/
│   │   └── forecast.xlsx
│   ├── Research/
│   │   └── market.pdf
│   └── Deliverables/
│
└── .apps_data/
    ├── mail/
    ├── chat/
    ├── calendar/
    └── fmp/

The filesystem/ directory contains ordinary task files such as spreadsheets, PDFs, Word documents, images, and CSVs. The .apps_data/ directory contains state for applications that are not naturally represented as standalone files. For example, a mail application might store:

{
  "messages": [
    {
      "from": "client@example.com",
      "subject": "Forecast question",
      "body": "Can you update the model?"
    }
  ]
}

A chat application might store channels and messages. A calendar application might store events.

These worlds are exactly what is stored in the Apex-Agents dataset, together with a bunch of tasks. The same world could be re-used across a different series of tasks.

The dataset then also contains the golden answers that we could use to evaluate the AI’s performance. For example, for the above example, the Apex-Agents dataset contains the following golden answer:

Complisure's updated 2030 revenue is expected to be $73,447,000 (low-end) 
to $78,830,000 (high-end).

The AI’s toolbox

How will the AI go about interacting with this world? It may want to read certain files, access certain information, change certain items, and so on. We need to re-create our tools into a format the AI can work with. This is where MCP servers come in.

The Archipelago agent environment includes MCP servers for applications such as: the filesystem, spreadsheets, PDFs, documents, presentations, mail, chat, calendar and code execution. These servers make the simulated workplace interactable through structured tools. For example, instead of an agent directly editing an .xlsx file itself, it can call a spreadsheet tool exposed by the spreadsheet MCP server.

An MCP server is just a program that exposes a set of tools in a structured way. At a high level, it does four things: declares which tools it provides, accepts structured tool calls, performs the requested operation and returns the result in a structured format.

For example, the spreadsheet MCP server might expose a tool such as:

sheets_server

Its server entry point lives at archipelago/mcp_servers/spreadsheets/mcp_servers/sheets_server/main.py. That file registers the spreadsheet functionality implemented under /mcp_servers/spreadsheets/mcp_servers/sheets_server/tools/. The actual Excel operations are handled with openpyxl.

And the way the agent sends off how exactly it wants to interact with the world happens through tool calls. This means the LLM must be good enough understanding the tool schema’s of the MCPs available and be able to output structured tool calls.

A tool call might look like this:

{
  "action": "read_tab",
  "file_path": "/5. Forecast model/forecast.xlsx",
  "tab_index": 0
}

When the server receives that request, it: resolves the requested file inside /filesystem, opens the workbook, reads the requested worksheet, converts the result into a structured response, returns that response to the caller. For an edit, the flow is similar: tool call, spreadsheet MCP server, open workbook, apply changes, save workbook, return result.

You can look through the MCP servers that Archipelago exposes to understand the setup. But for now, remember that these MCP servers are what will allow the AI agent to modify and interact with the world.

Wrapping this into an environment the AI can interact with

Now we need a way to turn this world into an environment the agent can interact with. Eventually, when we will do reinforcement learning, we will run many copies of these environments in parallel teaching the agent to get better at all of them as fast as possible, so the way we do this needs to be reproducible and robust.

Here we will use Archipelago as the runtime layer that starts the simulated workplace, loads the world data, and runs the services that make that world usable. The alternative is to use Harbor.

If we run,

cd archipelago/environment
docker compose up -d --build

it tells Docker to use the environment/Dockerfile and build the image, including Python, uv, LibreOffice, Chromium, system libraries, Archipelago dependencies, MCP server code and dependencies.

The DockerFile starts environment/runner/main.py which creates the FastAPI application and exposes endpoints such as:

GET  /health
POST /apps
POST /data/populate
POST /data/snapshot
POST /data/snapshot/s3
MCP  /mcp/

This application is the core Archipelago environment service. It handles the world state, application configuration, and snapshots. And remember the tools we discussed, that lives under arhcipelago/mcp_servers? Dockerfile copies these server implementations into the image and the MCP gateway can then launch whichever servers are needed for a particular environment, such as the filesystem, spreadsheet, mail, or calendar server.

Summarising the above. Once we have an idea of the information we need in the AI’s world and the specific task we want to teach it, we need to wrap this into a containerised Docker environment in which we can execute specific actions.

Archipelago is one framework to use for this. It is responsible for starting up this environment. Inside of the environment, we have a set of MCP servers, which will be exposed to the agent. It’s the main way the agent will interact with the world.

Replicating an AI’s interaction with the world

Remember the agent is just an LLM that outputs tool calls. In a bit, we will pass the starting prompt to an LLM, and see how it’ll go about solving our task.

But first, if we were to replicate how the LLM interacts with the world we created above, we can use the playground file in my GitHub and run:

uv run python playground.py --task-id task_d46f8183d88541c8ab7f2692aca28b5f

The command starts a fresh container, loads the world, configures the MCP

servers, prints the task prompt, and opens the shell. After this, you can interact and inspect the world in the same way the agent would.

Let’s go through how a human may go about this task, but using the commands that the agent would need to use.

The exact task in the selected Task ID above is:

The LATAM market customer count is expected to continue growing at its 2024-2025B CAGR. CompliSure could capture a quarter to a half of the two largest LatAm players’ latest share of customers if it expanded into the region. Forecast CompliSure’s 2030 revenue range.

We’re interacting with the environment through the MCP server. The agent does not automatically see the files. It first sees tool names and schemas, then uses those tools to discover the world. So a first step may be to expose the specific tools that are available to us.

mcp> tools

filesystem_server_list_files: List files and folders in a path
filesystem_server_search_files: Search for files matching a glob pattern in the given directory.
sheets_server_sheets: Spreadsheet operations: create, read, edit .xlsx and CSV files
sheets_server_sheets_schema: Get JSON schema for sheets input/output models.

To check the files that exist in our directory, we could do:

mcp> call filesystem_server_list_files {"path":"/"}

0. Project Briefing & Deliverables
4. Complisure internal data
2. Competitive analysis
1. TAM
3. Customer sentiment
6. Vertical SaaS deal case studies
5. Forecast model
7. Investment recommendation

mcp> call filesystem_server_list_files {"path":"/2. Competitive analysis"}
mcp> call filesystem_server_list_files {"path":"/5. Forecast model"}

We’d see amongst others, the following files:

/filesystem/
├── 2. Competitive analysis/
│   ├── Themes and Recommendations from Win_Loss Survey.docx
│   ├── Financial dataset_2016-25B expanded.xlsx
│   ├── 25 Sample Win_Loss Interview Responses.docx
│   ├── Feature dataset_competitors_2016-24.xlsx
│   ├── Competitive Landscape_competitors_all data.xlsx
│   ├── Geographic locations_competitors_2016-25B.xlsx
│   ├── Geographic locations_competitors_2016-25B expanded.xlsx
│   ├── Competitive analysis — CompliSure vs Vector Solutions, KPA, and Avetta.docx
│   └── Profile docs_all companies_Synthetic data/
│
└── 5. Forecast model/
    ├── 5yr_forecast_v5.xlsx
    ├── 4.5_Management_Forecast_5yr_PnL_v4.xlsx
    ├── CompliSure_5yr_Forecast_v2.xlsx
    ├── 4.5_Management_Forecast_5yr_PnL.xlsx
    ├── 4.5_Management_Forecast_Updated.xlsx
    └── CompliSure_5yr_Forecast.xlsx

The first folder contains a mixture of competitive-analysis spreadsheets, interview material, survey data, and company profiles. The second contains multiple versions of the management forecast and supporting assumptions.

To be precise: when we type the above command, the playground parses that into a tool name and arguments, then sends an MCP tools/call request to:

http://localhost:8080/mcp/

An LLM agent skips the human-readable shell command and directly emits a structured tool call, for example:

{
  "name": "filesystem_server_list_files",
  "arguments": {
    "path": "/"
  }
}

The agent runner forwards that call through MCP. The MCP Client packages that into an MCP request and sends it to the gateway and if that MCP server is not already running, the gateway can use its configuration to launch it as a subprocess and execute the specific call.

Which files would be relevant for our task? We can use the sheets_server_sheets command to understand what’s inside the files that we think are relevant.

mcp> call sheets_server_sheets {
  "action":"list_tabs",
  "file_path":"/2. Competitive analysis/Geographic locations_competitors_2016-25B expanded.xlsx"
}

Geographic locations_competitors_2016-25B expanded.xlsx
  Country data, index 0, 995 rows, 23 columns

mcp> call sheets_server_sheets {
  "action": "list_tabs",
  "file_path": "/2. Competitive analysis/Financial dataset_2016-25B expanded.xlsx"
}

Financial dataset_2016-25B expanded.xlsx
  All financials, index 0, 1000 rows, 18 columns

mcp> call sheets_server_sheets {
  "action": "list_tabs",
  "file_path": "/5. Forecast model/CompliSure_5yr_Forecast.xlsx"
}

CompliSure_5yr_Forecast.xlsx
  Sheet1, index 0, 7 rows, 8 columns

Let’s explore what’s in the geographic locations of competitors spreadsheet, filtering specifically on year 2025 and 2024 (remember, we needed to find the two largest competitors and the 2024 and 2025B LATAM totals to calculate the growth rate) and the LATAM market. We can use the filter_tab call from the server sheets MCP:

mcp> call sheets_server_sheets {
  "action":"filter_tab",
  "file_path":"/2. Competitive analysis/Geographic locations_competitors_2016-25B expanded.xlsx",
  "tab_index":0,
  "filters":[
    {"column":"Year","operator":"contains","value":"2025"},
    {"column":"Region","operator":"equals","value":"LATAM"}
  ],
   "match_all":true,"compact":true
}

Syncore: 16 Brazil + 29 Mexico = 45 LATAM customers
Velocity: 37 Brazil + 18 Mexico = 55 LATAM customers
UpSkill: 39 Brazil + 26 Mexico = 65 LATAM customers
TrainIQ: 33 Brazil + 37 Mexico = 70 LATAM customers

mcp> call sheets_server_sheets {
  "action":"filter_tab",
  "file_path":"/2. Competitive analysis/Geographic locations_competitors_2016-25B expanded.xlsx",
  "tab_index":0,
  "filters":[
    {"column":"Year","operator":"contains","value":"2024"},
    {"column":"Region","operator":"equals","value":"LATAM"}
  ],
  "match_all":true,"compact":true
}

Syncore: 15 Brazil + 28 Mexico = 43 LATAM customers
Velocity: 35 Brazil + 17 Mexico = 52 LATAM customers
UpSkill: 38 Brazil + 24 Mexico = 62 LATAM customers
TrainIQ: 32 Brazil + 35 Mexico = 67 LATAM customers

Great. We have discovered that the two largest players are TrainIQ with 70 customers and UpSkill with 65 customers. And that in total i) 2024 LATAM customers = 43 + 52 + 62 + 67 = 224, ii) 2025B LATAM customers = 45 + 55 + 65 + 70 = 235, so the 2024–2025B growth rate is: 235 / 224 - 1 = 4.91%.

By now, I presume you understand the flow. We are using the same tools as available to the agent, to figure out the information that exists in the Excel sheets. Sometimes, the agent may explore the incorrect tab, learn that information is stored somewhere else, redo the command and so on. Throughout, we are doing some interpretation of this information and reasoning on the calculations and next steps.

For the full walkthrough through the manual solution to this problem, see the here.

You can also play around with the other worlds, by launching the playground with another task ID.

Getting an LLM to interact with the world

Everything we did manually above, inspecting files, reading spreadsheets, choosing tools, and deciding what to do next, we now want an agent to do automatically.

To do that, we can construct a very small agent. At its core, the agent is just a loop that connects an LLM to the MCP tools exposed by the Archipelago environment.

On each iteration, the LLM receives:

  • the task it is trying to solve;
  • the tools currently available through MCP;
  • the results of any previous tool calls.

It then decides whether to call another tool or return a final answer. If it chooses a tool, the agent executes that call through MCP, adds the result back into the conversation, and asks the LLM what to do next.

To be precise, at startup the agent:

  • creates the conversation with the system prompt and task;
  • connects to the MCP gateway;
  • discovers the available MCP tools;
  • converts those tools into a format the LLM can use.

Then it enters a loop for a max_iters number of iterations:

task + tool schemas + previous results
            ↓
           LLM
            ↓
     tool call or answer

Specifically: the model response is parsed to check whether there is a tool call in it:

message = response.choices[0].message
tool_calls = message.get("tool_calls") or []

If the model requests a tool, the agent executes it through MCP:

result = await client.call_tool(name, arguments)

The result from the execution is then appended to the conversation and sent back to the model on the next iteration:

messages.append({
      "role": "tool",
      "tool_call_id": call.get("id"),
      "name": name,
      "content": result_text,
  })

response = await acompletion(
      model=model,
      messages=messages,
      tools=tools,
      tool_choice="auto",
      ...
  )

where the messages the model receives consist of the system prompt, the original task, previous assistant messages, previous tool results, all MCP tool schemas.

We need to extend this simple loop with two more things:

  • We need a way to decide when the agent is finished,
  • We need a summarisation prompt in case we reach the context window limit,
  • We need a way to handle incorrect tool calls.

Right now, the agent finishes once it reaches max_iters. But it may be done earlier than that. To fix this, we can expose a final_answer tool, and tell the agent to call this once it is done:

mcp_tools = await load_mcp_tools(...)
tools = list(mcp_tools) + [FINAL_ANSWER_TOOL]

The run is marked completed only when the model calls that tool: final_answer({”answer”:”...”}) or it stops when the maximum number of iterations is reached.

For the context window, the agent estimates conversation size using LiteLLM’s token_counter. When the estimate reaches that threshold, it asks the model for a compact progress summary containing: task, discovered files, tool results, calculations, decisions, unresolved questions, next action. Then it resumes with: system prompt, original task, progress summary.

Finally, the agent may output a tool call, the MCP server may attempt to execute this, but potentially this could throw an error as the tool call is wrong. This should not instantly shut down the agent, but instead, we’d like to append the tool call error to the context the agent will see:

 try:
      result = await client.call_tool(name, arguments)
      result_text = tool_result_text(result)
  except Exception as exc:
      result_text = (
          f"Tool call failed for {name}: "
          f"{type(exc).__name__}: {exc}. "
          "Inspect the tool schema and choose a compatible "
          "tool or arguments."
      )

In this way, even if the agent makes mistakes with the tool calls, it can recover from the failure using the error context.

My mini agent is implemented here, while Archipelago has its own agent implementation here. The scaffold that you place around the LLM is where a lot of work can go into, and it’s very customisable.

We can run the agent with either OpenAI or through a vLLM call which means that you’ll have to fire up a vLLM server with your model of choice.

The LLMs solution

Let’s use the vLLM setup. Note that you may want to create a separate environment for this:

uv venv --python 3.12 ~/vllm-env
uv pip install --python ~/vllm-env/bin/python vllm

We can choose to use a Qwen3-8B model with vLLM:

~/vllm-env/bin/vllm serve Qwen/Qwen3-8B \
    --host 0.0.0.0 \
    --port 8000 \
    --enable-auto-tool-choice \
    --tool-call-parser hermes

And then we can kick off the agent:

uv run python run_agent.py \
    --task-id task_d46f8183d88541c8ab7f2692aca28b5f \
    --model openai/Qwen/Qwen3-8B \
    --api-base http://localhost:8000/v1 \
    --api-key dummy \
    --max-steps 20 \
    --context-window-tokens 10000 \
    --output output/qwen3-8b/trajectory.json

The agent command will: reset/start Docker, load the selected APEX world, configure the MCP servers, connect to vLLM at localhost:8000, send all MCP tool schemas plus final_answer, execute the model’s tool calls and save the trajectory at for example examples/custom_agent/output/qwen3-8b/trajectory.json.

Let’s see the output. You can find the full trajectory here, but the model ended up doing:

  1. Inspect the filesystem The agent listed the root directory and found the relevantForecast model andCompetitive analysis folders.
  2. Read the baseline forecast It openedCompliSure_5yr_Forecast.xlsx and retrieved the 2030 baseline revenue:$73,230,550
  3. Inspect competitive-analysis files It listed the competitive-analysis directory, where the relevant geographic and financial workbooks were available.
  4. Choose the wrong workbook Instead of using the geographic-locations workbook to identify the largest LATAM competitors, it opened:Client share_competitors_2016-25B.xlsx The result was truncated, and the agent never retrieved the required LATAM customer counts or competitor financial data.
  5. Context compaction introduced unsupported assumptions After the context exceeded the token threshold, the trajectory was summarized. That summary incorrectly introduced claims about the largest LATAM competitors, their combined customer count, and an assumed contract size.
  6. Calculate from those assumptions Using the correct baseline revenue but unsupported competitive assumptions, the agent estimated additional revenue and produced:$73,358,000 to $73,485,500
  7. Submit the answer Thefinal_answer tool accepted the response, so the run ended with acompleted status.

Through manual observation, we can note that our agent completed the execution loop successfully, but not the underlying research task. The failure came from selecting the wrong source file and then relying on unsupported assumptions after context compaction.

Judging the quality

If we want to teach the agent LLM to get better at this task, having obtained a specific agent trajectory, we need to somehow grade this in an automated manner.

At its simplest, we can use another, larger LLM, to grade the final answer from our agent trajectory.

The judge receives the following prompt,

prompt = f"""Grade this answer against the task rubric. Return JSON only in the form {{
  "passed": true,
  "criteria": [
    {{
      "criterion": "...",
      "passed": true,
      "reason": "..."
    }}
  ]
  }}.

  Task:
  {task["prompt"]}

  Rubric:
  {json.dumps(task.get("rubric", []), indent=2)}

  Candidate answer:
  {answer}"""

The output from the judge is what will be used to further improve the agent’s ability to answer the question.

The logic of the judge itself can also be further improved to for e.g. also provide feedback on intermediate steps, such as penalise bad actions throughout the execution, validate that certain files were or were not modified, penalise too lengthy responses and so on. If you don’t want the model to escape a weakly protected sandbox like OpenAI did you may also penalise bad actions.

Let’s run our judge to see how it grades the trajectory:

uv run python src/grader.py   
  --task-id task_d46f8183d88541c8ab7f2692aca28b5f   
  --trajectory output/qwen3-8b/trajectory.json   
  --output output/qwen3-8b/grades.json

{
  "passed": false,
  "criteria": [
    {
      "criterion": "States the low end of Complisure's updated 2030 revenue is $73,447,000",
      "passed": false,
      "reason": "The candidate stated the low end as $73,358,000, which does not meet the specified figure."
    },
    {
      "criterion": "States the high end of Complisure's updated 2030 revenue is $78,830,000",
      "passed": false,
      "reason": "The candidate did not provide a high-end figure, stating instead a range that does not include the required high end."
    }
  ],
  "task_id": "task_d46f8183d88541c8ab7f2692aca28b5f",
  "mode": "llm"
}

Our judge correctly recognised that the trajectory from our agent was incorrect.

The quality of the judge is very important to get the model to improve. Our management consulting task is relatively open-ended. Other tasks such as maths or coding can have exact verification, e.g. in maths verify against a Lean environment (again, OpenAI’s Navier-Stokes agents were probably trained to do that), or with code against specific unit tests.

A bad manner of verification will never allow your model to properly improve.

Final step: running this all in an RL pipeline

This will be a fully separate blog post, but I just want to give you the outline of how all of the above connects to a reinforcement learning pipeline.

The current setup is an agent evaluation loop:

task prompt
  ↓
ReAct agent
  ↓
LLM chooses a tool
  ↓
MCP executes it in the Docker environment
  ↓
tool result returns to the model
  ↓
repeat
  ↓
final answer
  ↓
grader score

The agent receives the task, available MCP tools, and previous tool results. It keeps choosing either another tool call or a final answer. The full trajectory is then saved and graded.

The important point is that the model itself is not updated. RL adds a training step where the weights of the model will be updated based on the reward that was received from the judge. The specific strategy with which you update the algorithm is an active research domain. Lots of success was achieved with for example Group Relative Policy Optimization, as popularised in the DeepSeekMath paper, and many more variants of this have been introduced.

Then one RL update may look as such:

sample a world 
  ↓
run the agent in this world
  ↓
collect the rollout (also called trajectory)
  ↓
get the reward from the judge
  ↓
update the model/policy through something like GRPO or another RL algorithm
  ↓
do another rollout in another sampled world

The rollout is how the LLM goes about solving this task. It is exactly what we saw the LLM do above, e.g. something like:

reset environment of the world
  ↓
list files
  ↓
read spreadsheet
  ↓
filter data
  ↓
do some reasoning on the maths
  ↓
submit final answer
  ↓
receive reward

Typically, these RL steps are scaled across a huge number of GPUs, so that we can do many rollouts at the same time, collect rewards and update the agent model.

A framework for doing RL at scale efficiently is SkyRL. But this will be the next blogpost!

So, is your future life of freedom nearby?

Mercor shows that training Qwen3.5-397B-A17B on the white collar tasks we discussed above improved the AIs ability to solve the tasks end-to-end from 16.11% to 27.29%. In that same graph, they show that Grok-4.5 High scored the highest by solving 29%. At best, the above mechanism would allow to automate 30% or so of the tasks considered in the Apex-Agents benchmark.

We could expect this to increase, just like the METR graph has shown that the duration of tasks AI can work on independently is consistently going up, with top models now being able to work for 4 hours by themselves.

But there’s other caveats too.

Take the reinforcement learning itself. It’s not the case that the model can go from no abilities on these tasks to full ability. In practice, many frontier open reasoning models warm-start reinforcement learning with supervised fine-tuning on high-quality demonstrations, often including reasoning traces, before letting RL explore and improve the policy. Qwen3, Nemotron, and Tülu all follow variants of this recipe.

This means we need to obtain these expert demonstration traces, by having experts solve the task manually, in exactly the manner we did above: writing out the tool calls the model has to call, writing out the reasoning steps, perhaps showing how to recover from incorrect tool calls, and so on. Once again why we have the high valuations of the companies sourcing these experts.

Several other works (this, and this) discuss how the AI’s show limitations in creativity and judgment, and as the second work notes, they lack ‘the ability to check what they produced against what they found, revise when it does not hold up, and question whether the path they took was sound’, limiting the ability to consistently improve on the tasks at hand.

And even Anthropic’s report tells that ‘Claude is not operating fully autonomously for any measured subset of AI R&D work’. It’s peculiar to have both conversations of recursive self-improvement and civilisational risks posed by AI, while the reports on actual impact show very humble results for white-collar work.

And a final remark: Figure 2.5 here shows one of the highest AI autonomy rates for app and website development. And while some web developers I speak to say AI has made their work more enjoyable, others report a loss of enjoyment: no more time to look through the code, to actually understand the system, to create nice designs. AI has turned up the pressure of shipping features even faster.

Thanks for reading!

Anastasia

── more in #ai-agents 4 stories · sorted by recency
── more on @mercor 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/a-practical-guide-to…] indexed:0 read:24min 2026-09-30 · —