Thread-identity switcheroo for io_uring Io_uring maintainer Jens Axboe has posted an RFC patch set that adds a PF_IO_HANDOFF flag to the task_struct flags field, letting a thread that is about to block hand off its identity so the submitting thread never blocks. The change targets operations lacking asynchronous kernel support, such as the io_uring equivalents of fdatasync(), statx(), and some openat() paths, which current kernels push to a separate worker thread at the cost of a wakeup and context switch. Axboe's approach avoids reworking all kernel system-call paths to be non-blocking by proceeding with potentially blocking operations and handling the ones that actually block. Thread-identity switcheroo for io uring LWN subscriber-only content io uring subsystem https://man7.org/linux/man-pages/man7/io uring.7.html is all about asynchronous execution; applications count on it to not block — unless explicitly requested to. Within io uring, maintaining the "never blocks" guarantee has sometimes been a challenge, given that many paths in the kernel were never designed for asynchronous execution. This problem has been worked around, but at a significant cost to performance. Now, io uring maintainer Jens Axboe has posted an RFC patch set https://lwn.net/ml/all/20260911154148.644489-1-axboe@kernel.dk with a somewhat radical and potentially scary solution to the problem. A process in user space can submit one or more operations to io uring by describing them in submission-queue entries SQEs in the submission ring, then calling io uring enter https://man7.org/linux/man-pages/man2/io uring enter.2.html . The kernel will initiate processing on all of the entries that were placed into the ring. If possible, the kernel will execute them immediately within the io uring enter call. For example, if a read request can be satisfied with data that is already in the page cache and the user-space buffer is entirely resident within RAM, then the requested data can be copied immediately without blocking. In cases where blocking is required, though, things become more complicated. Many of the I/O paths within the kernel have asynchronous support built into them; they will carry a new request as far as it can go, and finish the job elsewhere within the kernel once the blocking operation has completed. There are other operations, though, that lack this support; these include the io uring equivalent of system calls like fdatasync https://man7.org/linux/man-pages/man2/fsync.2.html , statx https://man7.org/linux/man-pages/man2/statx.2.html , some openat https://man7.org/linux/man-pages/man2/openat.2.html paths the O NONBLOCK flag notwithstanding , and others. The io uring code must take extra care when implementing these operations, lest io uring enter block partway through processing a set of submitted operations. In current kernels, any operation that might cause io uring enter to block is handed off to a separate worker thread for execution. That allows the submitting thread to continue; the separate worker can block, if need be, without holding up anything else. This handoff is not free, though; the kernel must wake a waiting worker thread and perform a context switch, among other costs. If the operation does indeed block, those costs may not be significant in the end. In many cases, though, the operation can be completed without blocking. In such cases, the extra overhead becomes a significant part of the overall cost of executing the request. Since developers who turn to io uring are usually doing so in search of improved application performance, the cost of avoiding blocking that might not happen anyway hurts. The LWN kernel-source database https://lwn.net/ksdb/ is the definitive source of information about kernel releases. Try a one-month free trial subscription https://lwn.net/Promo/KSDB/claim for immediate access to LWN's kernel content and KSDB as well. One solution to the problem would be to rework all of the system-call paths in the kernel to be non-blocking. That has been done, in some cases, over the years, but it is not an easy or quick task. An alternative — the one that Axboe has chosen — is to proceed with a potentially blocking operation then handle cases that actually block without blocking the submitting thread. Detecting operations that do indeed block requires a change to the scheduler. A new flag PF IO HANDOFF is added to the flags field of the task struct structure https://elixir.bootlin.com/linux/v7.2.5/source/include/linux/sched.h L826 that represents a thread. If a thread is about to block for any reason and it has that flag set, the scheduler will make a call to a new function called io uring task sleeping . Hooking into the scheduler in this way allows io uring to be informed about blocking that happens anywhere in the kernel, without having to modify the actual code paths involved. Once io uring knows that an operation is going to block, it must do something about the situation. One option, in theory, would be to unwind whatever work had been done up to the blocking point, then to restart the operation in a worker thread. But, since this blocking can happen almost anywhere in the kernel, that is not really an option. As a general rule, once io uring has started a requested operation involving kernel code paths that are not designed to avoid blocking, it must see that operation through to the finish. Axboe's solution is "thread identity handoff". When the scheduler informs io uring about a thread that is about to block, io uring responds by selecting a worker thread from its thread pool. Rather than hand the ongoing work over to that thread which is not possible at this point , the code exchanges the identities of the two threads. The worker thread is made to look like the original submitting thread in every way, including its thread ID, signal-handling setup, and more; that thread then continues processing the submission ring before, eventually, returning to user space. The thread that returns from io uring enter has a different task struct than the one that made the call, but everything else hopefully looks the same. Meanwhile, the original thread, which was about to block executing an operation, takes on the worker thread's identity, then proceeds to block as usual. When it wakes, it will continue the operation through completion, then take its place in the worker-thread pool. The end result is that potentially blocking operations can be executed by the submitting thread and, if they can run without actually blocking, be handled entirely there. The cost of bringing in a worker thread is only paid if the operation really does block. It sounds simple enough, but this kind of identity exchange is fraught with potential land mines. Before the two threads involved can exchange their task struct structures, the kernel must make absolutely sure that nothing else in the kernel holds references to those structures. Otherwise, something will eventually be done with a reference to the wrong task struct , an outcome that will do nothing to reduce the strain on all of the people trying to keep up with the stream of kernel CVEs. That is a result that is deemed to be worth avoiding. Preventing it means being sure, before starting an operation that might block, that the submitting thread will be able to hand off its identity if the need arises. There is a long list found in the definitions of thread handoff allowed and thread handoff compatible in this patch https://lwn.net/ml/all/20260911154148.644489-2-axboe@kernel.dk of conditions that would prevent a handoff and require the operation to be executed in the old way. For example, if the thread is being traced with ptrace https://man7.org/linux/man-pages/man2/ptrace.2.html , then the tracer holds a reference to its task struct . By the same logic, if the thread in question is tracing any other tasks, those tasks hold references, so the thread cannot perform a handoff. Other conditions that will prevent a handoff include using perf events, having futex ownership tracked in the kernel, running under a realtime scheduler, holding a core scheduling https://lwn.net/Articles/861251/ cookie, performing a vfork https://man7.org/linux/man-pages/man2/vfork.2.html , and several others. This determination seems like the scariest, most fragile part of this series. The task struct is widely available, so it is hard to know that all of the possibilities for possible references in the kernel have been covered — before one even begins to worry about the addition of new references in the future by developers who are not thinking about io uring at all. The potential payoff is large, though. The cover letter includes a number of benchmark results. For some quick, non-blocking operations, the improvements can be huge; an fsync benchmark run on a tmpfs filesystem showed a nearly 700% improvement. Other improvements are more modest, and some of the tests that always block show regressions. The worst regressions tended to be with a higher queue depth — when there is a longer list of operations all being submitted at once. Immediately pushing each of those operations into a worker thread allows them to be worked on in parallel, while processing them up to the blocking point in the submitting thread serializes that work, slowing it down. Axboe said that he has ideas for addressing that problem, but he is unsure whether they are worth pursuing because, he said, applications that perform these operations tend not to have high queue depths to begin with. Axboe clearly does not expect to merge this series in the near future; he is more concerned with determining whether the overall approach has any chance of being viable. Comments have been limited so far. Peter Zijlstra pointed out https://lwn.net/ml/all/20260914114656.GC3500130@noisy.programming.kicks-ass.net/ that one of the other conditions blocking identity handoff — if the thread involved is running with a shadow stack — will prevent the use of the feature on most deployed systems; Axboe thinks https://lwn.net/ml/all/73837a51-e163-4b4e-8be2-4cdc044393b8@kernel.dk/ that the shadow stack can be moved with the rest of the thread's identity. Beyond that, it seems that developers are still mostly digesting this series that, as Gabriel Krisman Bertazi said https://lwn.net/ml/all/87tsnv1ynh.fsf@mailhost.krisman.be , is " really cool and seems like very dangerous thing ". If this series can convince developers that the "dangerous" part has been dealt with, it may eventually lead to significantly better io uring performance for a number of workloads. | Index entries for this article | | |---|---| | Kernel https://lwn.net/Kernel/Index | io uring https://lwn.net/Kernel/Index io uring |