# Thread-identity switcheroo for io_uring

> Source: <https://lwn.net/SubscriberLink/1094303/bf025f98cb71f941/>
> Published: 2026-09-27 23:48:14+00:00

# Thread-identity switcheroo for io_uring

## [LWN subscriber-only content]

[io_uring subsystem](https://man7.org/linux/man-pages/man7/io_uring.7.html)is all about asynchronous execution; applications count on it to not block — unless explicitly requested to. Within io_uring, maintaining the "never blocks" guarantee has sometimes been a challenge, given that many paths in the kernel were never designed for asynchronous execution. This problem has been worked around, but at a significant cost to performance. Now, io_uring maintainer Jens Axboe has posted

[an RFC patch set](https://lwn.net/ml/all/20260911154148.644489-1-axboe@kernel.dk)with a somewhat radical (and potentially scary) solution to the problem.

A process in user space can submit one or more operations to io_uring by
describing them in submission-queue entries (SQEs) in the submission ring,
then calling [`io_uring_enter()`](https://man7.org/linux/man-pages/man2/io_uring_enter.2.html).
The kernel will initiate processing on all of the entries that were placed into
the ring.  If possible, the kernel will execute them immediately within the
`io_uring_enter()` call.  For example, if a read request can be
satisfied with data that is already in the page cache and the user-space
buffer is entirely resident within RAM, then the requested data can be
copied immediately without blocking.

In cases where blocking *is* required, though, things become more
complicated.  Many of the I/O paths within the kernel have asynchronous
support built into them; they will carry a new request as far as it can go,
and finish the job elsewhere within the kernel once the blocking operation
has completed.  There are other operations, though, that lack this support;
these include the io_uring equivalent of system calls like [`fdatasync()`](https://man7.org/linux/man-pages/man2/fsync.2.html),
[`statx()`](https://man7.org/linux/man-pages/man2/statx.2.html),
some [`openat()`](https://man7.org/linux/man-pages/man2/openat.2.html)
paths (the `O_NONBLOCK` flag notwithstanding), and others.  The
io_uring code must take extra care when implementing these operations, lest
`io_uring_enter()` block partway through processing a set of
submitted operations.

In current kernels, any operation that *might* cause
`io_uring_enter()` to block is handed off to a separate worker
thread for execution.  That allows the submitting thread to continue; the
separate worker can block, if need be, without holding up anything else.
This handoff is not free, though; the kernel must wake a waiting worker
thread and perform a context switch, among other costs.  If the operation
does indeed block, those costs may not be significant in the end.  In many
cases, though, the operation can be completed without blocking.  In such
cases, the extra overhead becomes a significant part of the overall cost of
executing the request.  Since developers who turn to io_uring are usually
doing so in search of improved application performance, the cost of
avoiding blocking that might not happen anyway hurts.

The [LWN kernel-source database](https://lwn.net/ksdb/) is the definitive source of information about kernel releases.  [Try a one-month free trial subscription](https://lwn.net/Promo/KSDB/claim) for immediate access to LWN's kernel content and KSDB as well.

One solution to the problem would be to rework all of the system-call paths in the kernel to be non-blocking. That has been done, in some cases, over the years, but it is not an easy or quick task. An alternative — the one that Axboe has chosen — is to proceed with a potentially blocking operation then handle cases that actually block without blocking the submitting thread.

Detecting operations that do indeed block requires a change to the
scheduler.  A new flag (`PF_IO_HANDOFF`) is added to the
`flags` field of the [`task_struct`
structure](https://elixir.bootlin.com/linux/v7.2.5/source/include/linux/sched.h#L826) that represents a thread.  If a thread is about to block for
any reason and it has that flag set, the scheduler will make a call to a
new function called `io_uring_task_sleeping()`.  Hooking into the
scheduler in this way allows io_uring to be informed about blocking that
happens anywhere in the kernel, without having to modify the actual code
paths involved.

Once io_uring knows that an operation is going to block, it must do something about the situation. One option, in theory, would be to unwind whatever work had been done up to the blocking point, then to restart the operation in a worker thread. But, since this blocking can happen almost anywhere in the kernel, that is not really an option. As a general rule, once io_uring has started a requested operation involving kernel code paths that are not designed to avoid blocking, it must see that operation through to the finish.

Axboe's solution is "thread identity handoff".  When the scheduler informs
io_uring about a thread that is about to block, io_uring responds by
selecting a worker thread from its thread pool.  Rather than hand the
ongoing work over to that thread (which is not possible at this point), the
code exchanges the identities of the two threads.  The worker thread is
made to look like the original submitting thread in every way, including
its thread ID, signal-handling setup, and more; that thread then continues
processing the submission ring before, eventually, returning to user space.
The thread that returns from `io_uring_enter()` has a different
`task_struct` than the one that made the call, but everything else
(hopefully) looks the same.

Meanwhile, the original thread, which was about to block executing an operation, takes on the worker thread's identity, then proceeds to block as usual. When it wakes, it will continue the operation through completion, then take its place in the worker-thread pool. The end result is that potentially blocking operations can be executed by the submitting thread and, if they can run without actually blocking, be handled entirely there. The cost of bringing in a worker thread is only paid if the operation really does block.

It sounds simple enough, but this kind of identity exchange is fraught with
potential land mines.  Before the two threads involved can exchange their
`task_struct` structures, the kernel must make absolutely sure that
nothing else in the kernel holds references to those structures.
Otherwise, something will eventually be done with a reference to the wrong
`task_struct`, an outcome that will do nothing to reduce the strain
on all of the people trying to keep up with the stream of kernel CVEs.
That is a result that is deemed to be worth avoiding.

Preventing it means being sure, before starting an operation that might
block, that the submitting thread will be able to hand off its identity if
the need arises.  There is a long list (found in the definitions of
`thread_handoff_allowed()` and `thread_handoff_compatible()`
in [this
patch](https://lwn.net/ml/all/20260911154148.644489-2-axboe@kernel.dk)) of conditions that would prevent a handoff and require the
operation to be executed in the old way.  For example, if the
thread is being traced with [`ptrace()`](https://man7.org/linux/man-pages/man2/ptrace.2.html),
then the tracer holds a reference to its `task_struct`.  By the same
logic, if the thread in question is tracing any other tasks, those tasks
hold references, so the thread cannot perform a handoff.  Other conditions
that will prevent a handoff include using perf events, having futex
ownership tracked in the kernel, running under a realtime scheduler,
holding a [core scheduling](https://lwn.net/Articles/861251/) cookie,
performing a [`vfork()`](https://man7.org/linux/man-pages/man2/vfork.2.html),
and several others.  This determination seems like the scariest, most
fragile part of this series.  The `task_struct` is widely available,
so it is hard to know that all of the possibilities for possible references
in the kernel have been covered — before one even begins to worry about the
addition of new references in the future by developers who are not thinking
about io_uring at all.

The potential payoff is large, though.  The cover letter includes a number
of benchmark results.  For some quick, non-blocking operations, the
improvements can be huge; an `fsync()` benchmark run on a tmpfs
filesystem showed a nearly 700% improvement.  Other improvements are more
modest, and some of the tests that always block show regressions.  The
worst regressions tended to be with a higher queue depth — when there is a
longer list of operations all being submitted at once.  Immediately pushing
each of those operations into a worker thread allows them to be worked on
in parallel, while processing them up to the blocking point in the
submitting thread serializes that work, slowing it down.  Axboe said that
he has ideas for addressing that problem, but he is unsure whether they are
worth pursuing because, he said, applications that perform these operations
tend not to have high queue depths to begin with.

Axboe clearly does not expect to merge this series in the near future; he
is more concerned with determining whether the overall approach has any
chance of being viable.  Comments have been limited so far.  Peter Zijlstra
[pointed
out](https://lwn.net/ml/all/20260914114656.GC3500130@noisy.programming.kicks-ass.net/) that one of the other conditions blocking identity handoff — if the
thread involved is running with a shadow stack — will prevent the use of
the feature on most deployed systems; Axboe [thinks](https://lwn.net/ml/all/73837a51-e163-4b4e-8be2-4cdc044393b8@kernel.dk/)
that the shadow stack can be moved with the rest of the thread's identity.
Beyond that, it seems that developers are still mostly digesting this
series that, as Gabriel Krisman Bertazi [said](https://lwn.net/ml/all/87tsnv1ynh.fsf@mailhost.krisman.be), is "
really
cool and seems like very dangerous thing

".  If this series can convince
developers that the "dangerous" part has been dealt with, it may eventually
lead to significantly better io_uring performance for a number of workloads.

| Index entries for this article |  | 
|---|---|
| [Kernel](https://lwn.net/Kernel/Index) | [io_uring](https://lwn.net/Kernel/Index#io_uring) |
