GPT-6.1 Sol Async Tool Calling: Keep Agents Useful While a Human Decides OpenAI's GPT-6.1 Sol supports async tool calling, letting an application hand off a slow task or human question while the model continues independent work, but the API only removes the model-side wait and does not run jobs, store state, or track user responses. The model, positioned for complex coding, computer use, and professional work at lower cost than GPT-6 Astra, requires the Responses API and supports reasoning efforts from low through max but not none or minimal. Developers must build the orchestration layer themselves, including a task registry keyed by call_id and an application-generated task_handle, an idempotent result endpoint, an explicit wait policy, and tests for late or missing answers. A production pattern for the awkward middle of agent work: the agent needs one human answer, but the rest of the job should not freeze. Good agent UX does not turn one missing preference into a frozen workflow. Most AI agents fail the patience test in a very ordinary way. They find one missing detail — who a report is for, whether a deployment may proceed, which account to use — and stop. The user sees a spinner. The agent has other useful work it could do, but the tool loop treats every question like a roadblock. GPT-6.1 Sol makes a better pattern possible. Its async tool calling support lets an application hand off a slow task or a human question, then lets the model continue with independent work. The catch is important: the API removes the model-side wait; it does not run your jobs, store their state, chase a user response, or protect you from mixing up two similar calls. Your application still owns the hard part. This guide shows how to build that application layer. The running example is a report agent that needs an audience choice while it gathers facts and drafts an outline. The same design works for approvals, data exports, document generation, booking requests, and any tool that finishes on a different clock than the model. Async tools are not a shortcut around orchestration. They are a contract for keeping useful work moving while your system waits for something real. OpenAI positions GPT-6.1 Sol for complex coding, computer use, and professional work at a lower cost than GPT-6 Astra. It supports reasoning efforts from low through max, but not none or minimal. For tool calling, it requires the Responses API. Those details matter because an old, synchronous Chat Completions loop cannot simply gain this behavior through a new model name. The useful capability is simple to describe: mark an application-owned function or custom tool with async: true. The model may call it, keep reasoning, call other eligible tools, or return an independent answer. When the job or person eventually supplies the result, return it using the original call id. OpenAI’s async-tool guide also makes the boundary clear: hosted built-in tools are not async tools, async tools should use direct calling rather than Programmatic Tool Calling, and multi-agent mode should not combine them with parallel tool calls. That distinction creates a practical search gap. Broad explainers say “agents no longer need to wait.” Production teams still need the less glamorous pieces: a pending-task record, an idempotent result endpoint, an explicit wait policy, and tests for late or missing answers. Those pieces decide whether a quick demo becomes a trustworthy feature. A normal tool loop has one clock. The model requests a tool, the server runs it, the server returns its output, and the model continues. It works for a fast lookup. It becomes brittle when the work takes minutes or the “tool” is a person. An async workflow has two clocks: Do not pretend these clocks are synchronized. Instead, create a durable bridge between them. A small task registry is enough at first. Each row should hold the conversation ID, the original response ID, the model’s call id, an application-generated task handle, a task type, validated input, state, timestamps, and a safe result or error payload. Treat the call id as the receipt you must bring back to the model. Treat the task handle as the handle your application and UI use to find the work. The rule that prevents a surprisingly common bug: “Question displayed” is an event, not a tool result. A user-input tool is complete only when the user answers, declines, or times out with an explicit outcome. Keep the first tool intentionally boring. It should ask one answerable question, describe what can continue while it waits, and require a unique task handle. Avoid a vague ask user tool that can collect an essay, approval, secret, and account choice through one unbounded schema. Narrow tools are easier to render, audit, validate, and test. js const tools = { type: "function", name: "request audience async", async: true, strict: true, description: "Ask which audience should receive a report. Continue gathering facts and drafting a neutral outline while pending.", parameters: { type: "object", properties: { question: { type: "string" }, task handle: { type: "string" }, choices: { type: "array", items: { type: "string" }, minItems: 2, maxItems: 5 } }, required: "question", "task handle", "choices" , additionalProperties: false }}, { type: "function", name: "wait for tasks", strict: true, description: "Use only when the next step depends on a pending task result.", parameters: { type: "object", properties: { task handles: { type: "array", items: { type: "string" } } }, required: "task handles" , additionalProperties: false }} ; The tool description does real work. It tells the model to continue independent work and that waiting is a deliberate decision. Without that instruction, an agent may ask a good question and then immediately wait anyway. Use an application-generated handle such as audience 7f3c, not a value the model can accidentally reuse. Enforce uniqueness in the registry. A user reply that arrives after a similar second question should never attach to the wrong call. When a complete async function-call item arrives, parse and validate it. Then write the task record before you display the question or start a job. This order gives you recovery if a web process restarts halfway through rendering a modal or pushing a notification. The registry write should be idempotent by call id. If your stream reconnects or an event is delivered twice, a duplicate question is annoying; a duplicate charge or approval request can be much worse. Store the model’s original arguments as an immutable audit field and keep your normalized UI payload separately. For a human question, display the structured choices in your product. Do not return a fake function call output such as {"shown": true}. That would incorrectly tell the model the task is finished. The right state is pending. Let the response continue until the model either completes independent work or calls your wait tool because it truly needs the missing answer. The registry is the bridge between a model response and a result that arrives later. When the user selects “engineering leaders,” your application reads the pending row, checks that the actor is allowed to answer it, marks the task completed, and continues from the latest response. The continuation includes the original call id and the real result. It also includes the tools and instructions again so the model has the same operating contract. js async function submitAudienceAnswer taskHandle, audience { const task = await tasks.completeOnce { taskHandle, output: { task handle: taskHandle, audience } } ; if task return; // already completed or no longer valid return openai.responses.create { model: "gpt-6.1-sol", previous response id: task.latest response id, input: { type: "function call output", call id: task.call id, output: JSON.stringify task.output } , tools, instructions: REPORT AGENT INSTRUCTIONS } ;} Store the latest response ID whenever the model produces a new response. Do not assume the response that issued the question is still the correct continuation point after the model has performed independent work. If several pending tasks can resolve in any order, serialize continuation work per conversation or use a clearly designed event processor. Otherwise, two valid results can race and one continuation can silently lose context. For jobs rather than people, apply the same pattern. A data export worker posts a compact, validated result to your task service. Your service deduplicates it, verifies the job belongs to the pending handle, and feeds it back on the original call. The model gets a fact, not a raw worker log. Keep oversized artifacts in your own storage and return a reference or a deliberately reduced summary. Async tool calling is easy to misuse as a badge of sophistication. A one-second account lookup is usually clearer as a normal direct tool call. A predictable batch of fetches and joins may fit Programmatic Tool Calling better. OpenAI describes that mode as generated JavaScript for bounded, predictable orchestration; it is a different control surface from an application-owned task that returns later. Async calling earns its complexity when three conditions hold: A report agent fits. It can gather sources, identify assumptions, and prepare two outline variants while a manager chooses the audience. An approval-sensitive deployment often does not fit until you draw a hard boundary: it may prepare a plan and evidence while waiting, but it must not perform the deployment or generate a success message before the approval result arrives. Async workflows need stopping rules. Otherwise, the model may keep working long after the remaining task has become essential, or it may repeatedly ask the user in new words. Put the policy in the system or developer instructions.