cd /news/ai-agents/gpt-6-1-sol-async-tool-calling-keep-… · home › topics › ai-agents › article
[ARTICLE · art-147037] src=pub.towardsai.net ↗ pub= topic=ai-agents verified=true sentiment=· neutral

GPT-6.1 Sol Async Tool Calling: Keep Agents Useful While a Human Decides

OpenAI's GPT-6.1 Sol supports async tool calling, letting an application hand off a slow task or human question while the model continues independent work, but the API only removes the model-side wait and does not run jobs, store state, or track user responses. The model, positioned for complex coding, computer use, and professional work at lower cost than GPT-6 Astra, requires the Responses API and supports reasoning efforts from low through max but not none or minimal. Developers must build the orchestration layer themselves, including a task registry keyed by call_id and an application-generated task_handle, an idempotent result endpoint, an explicit wait policy, and tests for late or missing answers.

by read12 min views2 publishedOct 7, 2026

A production pattern for the awkward middle of agent work: the agent needs one human answer, but the rest of the job should not freeze.

Good agent UX does not turn one missing preference into a frozen workflow.

Most AI agents fail the patience test in a very ordinary way. They find one missing detail — who a report is for, whether a deployment may proceed, which account to use — and stop. The user sees a spinner. The agent has other useful work it could do, but the tool loop treats every question like a roadblock.

GPT-6.1 Sol makes a better pattern possible. Its async tool calling support lets an application hand off a slow task or a human question, then lets the model continue with independent work. The catch is important: the API removes the model-side wait; it does not run your jobs, store their state, chase a user response, or protect you from mixing up two similar calls. Your application still owns the hard part.

This guide shows how to build that application layer. The running example is a report agent that needs an audience choice while it gathers facts and drafts an outline. The same design works for approvals, data exports, document generation, booking requests, and any tool that finishes on a different clock than the model.

Async tools are not a shortcut around orchestration. They are a contract for keeping useful work moving while your system waits for something real.

OpenAI positions GPT-6.1 Sol for complex coding, computer use, and professional work at a lower cost than GPT-6 Astra. It supports reasoning efforts from low through max, but not none or minimal. For tool calling, it requires the Responses API. Those details matter because an old, synchronous Chat Completions loop cannot simply gain this behavior through a new model name.

The useful capability is simple to describe: mark an application-owned function or custom tool with async: true. The model may call it, keep reasoning, call other eligible tools, or return an independent answer. When the job or person eventually supplies the result, return it using the original call_id. OpenAI’s async-tool guide also makes the boundary clear: hosted built-in tools are not async tools, async tools should use direct calling rather than Programmatic Tool Calling, and multi-agent mode should not combine them with parallel tool calls.

That distinction creates a practical search gap. Broad explainers say “agents no longer need to wait.” Production teams still need the less glamorous pieces: a pending-task record, an idempotent result endpoint, an explicit wait policy, and tests for late or missing answers. Those pieces decide whether a quick demo becomes a trustworthy feature.

A normal tool loop has one clock. The model requests a tool, the server runs it, the server returns its output, and the model continues. It works for a fast lookup. It becomes brittle when the work takes minutes or the “tool” is a person.

An async workflow has two clocks:

Do not pretend these clocks are synchronized. Instead, create a durable bridge between them. A small task registry is enough at first. Each row should hold the conversation ID, the original response ID, the model’s call_id, an application-generated task_handle, a task type, validated input, state, timestamps, and a safe result or error payload. Treat the call_id as the receipt you must bring back to the model. Treat the task_handle as the handle your application and UI use to find the work.

The rule that prevents a surprisingly common bug: “Question displayed” is an event, not a tool result. A user-input tool is complete only when the user answers, declines, or times out with an explicit outcome.

Keep the first tool intentionally boring. It should ask one answerable question, describe what can continue while it waits, and require a unique task handle. Avoid a vague ask_user tool that can collect an essay, approval, secret, and account choice through one unbounded schema. Narrow tools are easier to render, audit, validate, and test.

const tools = [{  type: "function", name: "request_audience_async", async: true, strict: true,  description: "Ask which audience should receive a report. Continue gathering facts and drafting a neutral outline while pending.",  parameters: { type: "object", properties: {    question: { type: "string" }, task_handle: { type: "string" },    choices: { type: "array", items: { type: "string" }, minItems: 2, maxItems: 5 }  }, required: ["question", "task_handle", "choices"], additionalProperties: false }}, {  type: "function", name: "wait_for_tasks", strict: true,  description: "Use only when the next step depends on a pending task result.",  parameters: { type: "object", properties: { task_handles: { type: "array", items: { type: "string" } } }, required: ["task_handles"], additionalProperties: false }}];

The tool description does real work. It tells the model to continue independent work and that waiting is a deliberate decision. Without that instruction, an agent may ask a good question and then immediately wait anyway. Use an application-generated handle such as audience_7f3c, not a value the model can accidentally reuse. Enforce uniqueness in the registry. A user reply that arrives after a similar second question should never attach to the wrong call.

When a complete async function-call item arrives, parse and validate it. Then write the task record before you display the question or start a job. This order gives you recovery if a web process restarts halfway through rendering a modal or pushing a notification.

The registry write should be idempotent by call_id. If your stream reconnects or an event is delivered twice, a duplicate question is annoying; a duplicate charge or approval request can be much worse. Store the model’s original arguments as an immutable audit field and keep your normalized UI payload separately.

For a human question, display the structured choices in your product. Do not return a fake function_call_output such as {"shown": true}. That would incorrectly tell the model the task is finished. The right state is pending. Let the response continue until the model either completes independent work or calls your wait tool because it truly needs the missing answer.

The registry is the bridge between a model response and a result that arrives later.

When the user selects “engineering leaders,” your application reads the pending row, checks that the actor is allowed to answer it, marks the task completed, and continues from the latest response. The continuation includes the original call_id and the real result. It also includes the tools and instructions again so the model has the same operating contract.

async function submitAudienceAnswer(taskHandle, audience) {  const task = await tasks.completeOnce({ taskHandle, output: { task_handle: taskHandle, audience } });  if (!task) return; // already completed or no longer valid  return openai.responses.create({    model: "gpt-6.1-sol", previous_response_id: task.latest_response_id,    input: [{ type: "function_call_output", call_id: task.call_id, output: JSON.stringify(task.output) }],    tools, instructions: REPORT_AGENT_INSTRUCTIONS  });}

Store the latest response ID whenever the model produces a new response. Do not assume the response that issued the question is still the correct continuation point after the model has performed independent work. If several pending tasks can resolve in any order, serialize continuation work per conversation or use a clearly designed event processor. Otherwise, two valid results can race and one continuation can silently lose context.

For jobs rather than people, apply the same pattern. A data export worker posts a compact, validated result to your task service. Your service deduplicates it, verifies the job belongs to the pending handle, and feeds it back on the original call. The model gets a fact, not a raw worker log. Keep oversized artifacts in your own storage and return a reference or a deliberately reduced summary.

Async tool calling is easy to misuse as a badge of sophistication. A one-second account lookup is usually clearer as a normal direct tool call. A predictable batch of fetches and joins may fit Programmatic Tool Calling better. OpenAI describes that mode as generated JavaScript for bounded, predictable orchestration; it is a different control surface from an application-owned task that returns later.

Async calling earns its complexity when three conditions hold:

A report agent fits. It can gather sources, identify assumptions, and prepare two outline variants while a manager chooses the audience. An approval-sensitive deployment often does not fit until you draw a hard boundary: it may prepare a plan and evidence while waiting, but it must not perform the deployment or generate a success message before the approval result arrives.

Async workflows need stopping rules. Otherwise, the model may keep working long after the remaining task has become essential, or it may repeatedly ask the user in new words. Put the policy in the system or developer instructions.

<async_task_policy>When an async task is pending, continue only work that remains correct without its result. Do not infer, invent, or acknowledge the pending result. Call wait_for_tasks when the next irreversible action, final recommendation, or final response depends on it. Ask each decision once. If a task returns an error, timeout, decline, or cancellation, explain the limitation and offer a safe next step. Never perform approval-sensitive actions while approval is pending.</async_task_policy>

The word “irreversible” is useful. It captures sending an email, changing a record, publishing a report, reserving inventory, and claiming a decision has been made. The model can create a draft; it cannot represent a draft as an approved outcome.

Every pending task needs a deadline. When it expires, return a structured outcome such as {"status":"timed_out"} on the original call rather than deleting the task. The model can then ask a shorter follow-up, offer a default only if your policy allows it, or finish the parts that are still valid. A vanished task creates an unexplained loop.

Users double-click. Webhooks retry. A background worker completes, then crashes before it records success. Design the completion path so only the first valid result changes a pending task to completed. Later deliveries should return the recorded completion without creating a second continuation.

A task handle is not authorization. Check that the person answering the question belongs to the right workspace and conversation. For approvals, record who answered, which policy version applied, and the exact prompt shown. Treat a link containing a task handle as a locator, never as proof of permission.

Returning a thirty-page vendor payload tempts the model to miss the key field or restate unsafe content. Define a result schema. For the audience tool, return a single allowed choice. For an export, return status, artifact reference, row count, and warnings. Let your worker validate the result before the model sees it.

The model’s response and the task system need separate, observable lifecycles.

A demo proves that one answer can arrive. A production test plan proves that every awkward answer leaves the conversation in a coherent state. Build a fixture conversation with an audience question, a slow report lookup, and work that is safe to complete while both are pending.

Then test these cases:

Assertions should cover more than final prose. Check the registry state, call-ID linkage, number of user prompts, event order, identity checks, side effects, and whether the final response accurately says what remains pending. For GPT-6.1 Sol, also record reasoning effort, latency, input and output tokens, cached tokens, tool calls, and final task success. A lower token count is not an improvement if the agent waited too soon or skipped a necessary question.

Async tools have a human promise: the product should keep feeling alive. Measure that promise directly. Track time to first useful progress after an async call, the share of conversations where independent work was completed before waiting, task completion rate, duplicate completions prevented, timeout rate, and user abandonment while a question is pending.

Pair those numbers with quality checks. Did the agent ask only decision-changing questions? Did it expose a draft as final? Did it make a protected write while approval was pending? Did the final response connect the returned answer to the work it changed? These are better signals than a generic “async throughput” number.

Start with one read-only or draft-only workflow. Use one async question tool with a handful of choices. Add the registry, the timeout path, and the replay tests before you add more task types. Once the flow is observable, introduce a slow job tool. Only then consider approval-like tasks, and keep the actual write behind a direct, auditable boundary.

GPT-6.1 Sol is a strong fit when the work is complex enough to benefit from reasoning and tools, yet cost still matters. But the model choice is not the whole design. The durable task record, original call_id, explicit wait rule, and boring idempotency checks are what turn a clever async demo into an agent people can rely on.

It is a Responses API pattern where an application-owned function or custom tool is marked async: true. The model can continue independent work while the application waits for the tool result, then receives that result later on the original call ID.

No. Your application runs the job, stores its state, handles retries and authorization, and returns the result. Async calling prevents the model from having to block on the result.

For tool calling, GPT-6.1 Sol requires the Responses API. Chat Completions can be used without tools, but it is not the right integration path for this workflow.

No. Use it when the result is truly delayed and the model has useful work that remains correct without it. A short lookup is usually simpler as a normal direct call. A predictable batch may fit Programmatic Tool Calling better.

Expire the task and return an explicit timeout or no-answer result on the original call ID. The model can then explain what is blocked, continue safe work, or ask a focused follow-up according to your policy.

Yes, but only with a strong boundary. The model can prepare evidence while approval is pending, but the protected action must wait for a verified approval result and should remain directly auditable.

Technical behavior in this article is based on official OpenAI documentation for GPT-6.1 Sol, async tool calling, function calling, and Programmatic Tool Calling.

GPT-6.1 Sol Async Tool Calling: Keep Agents Useful While a Human Decides was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.

── more in #ai-agents 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/gpt-6-1-sol-async-to…] indexed:0 read:12min 2026-10-07 · —