{"slug": "thread-identity-switcheroo-for-io-uring", "title": "Thread-identity switcheroo for io_uring", "summary": "Io_uring maintainer Jens Axboe has posted an RFC patch set that adds a PF_IO_HANDOFF flag to the task_struct flags field, letting a thread that is about to block hand off its identity so the submitting thread never blocks. The change targets operations lacking asynchronous kernel support, such as the io_uring equivalents of fdatasync(), statx(), and some openat() paths, which current kernels push to a separate worker thread at the cost of a wakeup and context switch. Axboe's approach avoids reworking all kernel system-call paths to be non-blocking by proceeding with potentially blocking operations and handling the ones that actually block.", "body_md": "# Thread-identity switcheroo for io_uring\n\n## [LWN subscriber-only content]\n\n[io_uring subsystem](https://man7.org/linux/man-pages/man7/io_uring.7.html)is all about asynchronous execution; applications count on it to not block — unless explicitly requested to. Within io_uring, maintaining the \"never blocks\" guarantee has sometimes been a challenge, given that many paths in the kernel were never designed for asynchronous execution. This problem has been worked around, but at a significant cost to performance. Now, io_uring maintainer Jens Axboe has posted\n\n[an RFC patch set](https://lwn.net/ml/all/20260911154148.644489-1-axboe@kernel.dk)with a somewhat radical (and potentially scary) solution to the problem.\n\nA process in user space can submit one or more operations to io_uring by\ndescribing them in submission-queue entries (SQEs) in the submission ring,\nthen calling [`io_uring_enter()`](https://man7.org/linux/man-pages/man2/io_uring_enter.2.html).\nThe kernel will initiate processing on all of the entries that were placed into\nthe ring.  If possible, the kernel will execute them immediately within the\n`io_uring_enter()` call.  For example, if a read request can be\nsatisfied with data that is already in the page cache and the user-space\nbuffer is entirely resident within RAM, then the requested data can be\ncopied immediately without blocking.\n\nIn cases where blocking *is* required, though, things become more\ncomplicated.  Many of the I/O paths within the kernel have asynchronous\nsupport built into them; they will carry a new request as far as it can go,\nand finish the job elsewhere within the kernel once the blocking operation\nhas completed.  There are other operations, though, that lack this support;\nthese include the io_uring equivalent of system calls like [`fdatasync()`](https://man7.org/linux/man-pages/man2/fsync.2.html),\n[`statx()`](https://man7.org/linux/man-pages/man2/statx.2.html),\nsome [`openat()`](https://man7.org/linux/man-pages/man2/openat.2.html)\npaths (the `O_NONBLOCK` flag notwithstanding), and others.  The\nio_uring code must take extra care when implementing these operations, lest\n`io_uring_enter()` block partway through processing a set of\nsubmitted operations.\n\nIn current kernels, any operation that *might* cause\n`io_uring_enter()` to block is handed off to a separate worker\nthread for execution.  That allows the submitting thread to continue; the\nseparate worker can block, if need be, without holding up anything else.\nThis handoff is not free, though; the kernel must wake a waiting worker\nthread and perform a context switch, among other costs.  If the operation\ndoes indeed block, those costs may not be significant in the end.  In many\ncases, though, the operation can be completed without blocking.  In such\ncases, the extra overhead becomes a significant part of the overall cost of\nexecuting the request.  Since developers who turn to io_uring are usually\ndoing so in search of improved application performance, the cost of\navoiding blocking that might not happen anyway hurts.\n\nThe [LWN kernel-source database](https://lwn.net/ksdb/) is the definitive source of information about kernel releases.  [Try a one-month free trial subscription](https://lwn.net/Promo/KSDB/claim) for immediate access to LWN's kernel content and KSDB as well.\n\nOne solution to the problem would be to rework all of the system-call paths in the kernel to be non-blocking. That has been done, in some cases, over the years, but it is not an easy or quick task. An alternative — the one that Axboe has chosen — is to proceed with a potentially blocking operation then handle cases that actually block without blocking the submitting thread.\n\nDetecting operations that do indeed block requires a change to the\nscheduler.  A new flag (`PF_IO_HANDOFF`) is added to the\n`flags` field of the [`task_struct`\nstructure](https://elixir.bootlin.com/linux/v7.2.5/source/include/linux/sched.h#L826) that represents a thread.  If a thread is about to block for\nany reason and it has that flag set, the scheduler will make a call to a\nnew function called `io_uring_task_sleeping()`.  Hooking into the\nscheduler in this way allows io_uring to be informed about blocking that\nhappens anywhere in the kernel, without having to modify the actual code\npaths involved.\n\nOnce io_uring knows that an operation is going to block, it must do something about the situation. One option, in theory, would be to unwind whatever work had been done up to the blocking point, then to restart the operation in a worker thread. But, since this blocking can happen almost anywhere in the kernel, that is not really an option. As a general rule, once io_uring has started a requested operation involving kernel code paths that are not designed to avoid blocking, it must see that operation through to the finish.\n\nAxboe's solution is \"thread identity handoff\".  When the scheduler informs\nio_uring about a thread that is about to block, io_uring responds by\nselecting a worker thread from its thread pool.  Rather than hand the\nongoing work over to that thread (which is not possible at this point), the\ncode exchanges the identities of the two threads.  The worker thread is\nmade to look like the original submitting thread in every way, including\nits thread ID, signal-handling setup, and more; that thread then continues\nprocessing the submission ring before, eventually, returning to user space.\nThe thread that returns from `io_uring_enter()` has a different\n`task_struct` than the one that made the call, but everything else\n(hopefully) looks the same.\n\nMeanwhile, the original thread, which was about to block executing an operation, takes on the worker thread's identity, then proceeds to block as usual. When it wakes, it will continue the operation through completion, then take its place in the worker-thread pool. The end result is that potentially blocking operations can be executed by the submitting thread and, if they can run without actually blocking, be handled entirely there. The cost of bringing in a worker thread is only paid if the operation really does block.\n\nIt sounds simple enough, but this kind of identity exchange is fraught with\npotential land mines.  Before the two threads involved can exchange their\n`task_struct` structures, the kernel must make absolutely sure that\nnothing else in the kernel holds references to those structures.\nOtherwise, something will eventually be done with a reference to the wrong\n`task_struct`, an outcome that will do nothing to reduce the strain\non all of the people trying to keep up with the stream of kernel CVEs.\nThat is a result that is deemed to be worth avoiding.\n\nPreventing it means being sure, before starting an operation that might\nblock, that the submitting thread will be able to hand off its identity if\nthe need arises.  There is a long list (found in the definitions of\n`thread_handoff_allowed()` and `thread_handoff_compatible()`\nin [this\npatch](https://lwn.net/ml/all/20260911154148.644489-2-axboe@kernel.dk)) of conditions that would prevent a handoff and require the\noperation to be executed in the old way.  For example, if the\nthread is being traced with [`ptrace()`](https://man7.org/linux/man-pages/man2/ptrace.2.html),\nthen the tracer holds a reference to its `task_struct`.  By the same\nlogic, if the thread in question is tracing any other tasks, those tasks\nhold references, so the thread cannot perform a handoff.  Other conditions\nthat will prevent a handoff include using perf events, having futex\nownership tracked in the kernel, running under a realtime scheduler,\nholding a [core scheduling](https://lwn.net/Articles/861251/) cookie,\nperforming a [`vfork()`](https://man7.org/linux/man-pages/man2/vfork.2.html),\nand several others.  This determination seems like the scariest, most\nfragile part of this series.  The `task_struct` is widely available,\nso it is hard to know that all of the possibilities for possible references\nin the kernel have been covered — before one even begins to worry about the\naddition of new references in the future by developers who are not thinking\nabout io_uring at all.\n\nThe potential payoff is large, though.  The cover letter includes a number\nof benchmark results.  For some quick, non-blocking operations, the\nimprovements can be huge; an `fsync()` benchmark run on a tmpfs\nfilesystem showed a nearly 700% improvement.  Other improvements are more\nmodest, and some of the tests that always block show regressions.  The\nworst regressions tended to be with a higher queue depth — when there is a\nlonger list of operations all being submitted at once.  Immediately pushing\neach of those operations into a worker thread allows them to be worked on\nin parallel, while processing them up to the blocking point in the\nsubmitting thread serializes that work, slowing it down.  Axboe said that\nhe has ideas for addressing that problem, but he is unsure whether they are\nworth pursuing because, he said, applications that perform these operations\ntend not to have high queue depths to begin with.\n\nAxboe clearly does not expect to merge this series in the near future; he\nis more concerned with determining whether the overall approach has any\nchance of being viable.  Comments have been limited so far.  Peter Zijlstra\n[pointed\nout](https://lwn.net/ml/all/20260914114656.GC3500130@noisy.programming.kicks-ass.net/) that one of the other conditions blocking identity handoff — if the\nthread involved is running with a shadow stack — will prevent the use of\nthe feature on most deployed systems; Axboe [thinks](https://lwn.net/ml/all/73837a51-e163-4b4e-8be2-4cdc044393b8@kernel.dk/)\nthat the shadow stack can be moved with the rest of the thread's identity.\nBeyond that, it seems that developers are still mostly digesting this\nseries that, as Gabriel Krisman Bertazi [said](https://lwn.net/ml/all/87tsnv1ynh.fsf@mailhost.krisman.be), is \"\nreally\ncool and seems like very dangerous thing\n\n\".  If this series can convince\ndevelopers that the \"dangerous\" part has been dealt with, it may eventually\nlead to significantly better io_uring performance for a number of workloads.\n\n| Index entries for this article |  | \n|---|---|\n| [Kernel](https://lwn.net/Kernel/Index) | [io_uring](https://lwn.net/Kernel/Index#io_uring) |", "url": "https://wpnews.pro/news/thread-identity-switcheroo-for-io-uring", "canonical_source": "https://lwn.net/SubscriberLink/1094303/bf025f98cb71f941/", "published_at": "2026-09-27 23:48:14+00:00", "updated_at": "2026-09-28 00:01:13.917317+00:00", "lang": "en", "topics": ["ai-infrastructure", "developer-tools"], "entities": ["Jens Axboe", "io_uring", "Linux kernel", "task_struct", "PF_IO_HANDOFF", "LWN"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/thread-identity-switcheroo-for-io-uring", "markdown": "https://wpnews.pro/news/thread-identity-switcheroo-for-io-uring.md", "text": "https://wpnews.pro/news/thread-identity-switcheroo-for-io-uring.txt", "jsonld": "https://wpnews.pro/news/thread-identity-switcheroo-for-io-uring.jsonld"}}