{"slug": "data-loss-in-ttrpc-rust", "title": "Data loss in ttrpc-rust", "summary": "A data-loss bug in containerd's ttrpc-rust protocol implementation caused the final log frame of a stream to be silently dropped, traced to a per-frame tokio::spawn in src/asynchronous/client.rs at commit f31f592 that let a FLAG_REMOTE_CLOSED frame remove a stream from the client's streams map before the preceding DATA frame's task looked it up. The author's fix removes the spawn so the reader task handles each frame in read order, but a reviewer noted the naive fix lets a full 100-message bounded mpsc channel block the reader and stall every other call on the connection, as reported in issue #311.", "body_md": "*The interactive version of this post, with animations, is at [notes.shvbsle.in/ttrpc-data-loss](https://notes.shvbsle.in/ttrpc-data-loss/).*\n\nBefore the AI-accelerated programming age, it was expected from a software stack that all of its stable components are written in one uniform language. In the case of containers, the ecosystem (containerd, shim, runc) was written in Golang. But in this new age, armed with LLMs, any new language can make lateral entry in this ecosystem and quickly build stable runtimes. I've been trying to introduce more Rust to this ecosystem.\n\nOne day I was watching logs from a pod that was running on a container runtime written in Rust and noticed that occasionally the log stream just abruptly ended. I deployed another pod that prints numbers from 1 to 100, and the logs sometimes abruptly ended on 99, but never on 97 or 98. I kept thinking that it was a bug in the runtime, but it turned out that the bug was hidden much deeper in the stack. It was in the implementation of the protocol<sup>[1](#fn-1)</sup>. This page contains my notes on the internals of ttrpc and ways in which this bug manifests.\n\n*Interactive: [step through real ttrpc frames byte by byte](https://notes.shvbsle.in/ttrpc-data-loss/#overview-of-the-ttrpc-protocol).*\n\n`rt-multi-thread`).` 05`) removes the stream ID from the client's `streams` map.\n*Interactive: [watch the race drop frame 100](https://notes.shvbsle.in/ttrpc-data-loss/#feel-the-bug).*\n\nIn code, it was this:\n\n``` js\nasync fn handle_msg(&self, msg: GenMessage) {\n    let req_map = self.streams.clone();\n    tokio::spawn(async move {\n        if let Some(resp_tx) = get_resp_tx(req_map, &msg.header).await {\n            resp_tx\n                .send(Ok(msg))\n                .await\n                .unwrap_or_else(|_e| error!(\"The request has returned\"));\n        }\n    });\n}\n```\n\n[src/asynchronous/client.rs @ f31f592](https://github.com/containerd/ttrpc-rust/blob/f31f5925749bba2616bc53942dd83877a3f5b532/src/asynchronous/client.rs#L327-L337) · [issue #311](https://github.com/containerd/ttrpc-rust/issues/311)\n\nStop spawning: the reader task handles each frame itself, in the order it read them.\n\n```\n async fn handle_msg(&self, msg: GenMessage) {\n+    // Do not `tokio::spawn` per frame: a `FLAG_REMOTE_CLOSED` frame could\n+    // then `remove` a stream from `req_map` before the preceding DATA\n+    // frame's task looked it up, silently dropping the final payload.\n+    // The read loop already awaits this per frame, so inline is correct.\n     let req_map = self.streams.clone();\n-    tokio::spawn(async move {\n-        if let Some(resp_tx) = get_resp_tx(req_map, &msg.header).await {\n-            resp_tx\n-                .send(Ok(msg))\n-                .await\n-                .unwrap_or_else(|_e| error!(\"The request has returned\"));\n-        }\n-    });\n+    if let Some(resp_tx) = get_resp_tx(req_map, &msg.header).await {\n+        resp_tx\n+            .send(Ok(msg))\n+            .await\n+            .unwrap_or_else(|_e| error!(\"The request has returned\"));\n+    }\n }\n```\n\n*Interactive: [the reader handling every frame in order](https://notes.shvbsle.in/ttrpc-data-loss/#naive-fix).*\n\nBut this naive solution was NOT perfect. My \"reality has a surprising amount of detail\"<sup>[4](#fn-4)</sup> moment happened when the reviewer<sup>[5](#fn-5)</sup> made me aware of a different limitation of this fix.\n\nThere is another construct that we must think of: the bounded mpsc channel. Every call's frames reach its caller through a tokio mpsc channel that holds 100 messages, and `send().await` waits while it is full.\n\n```\nlet (tx, rx): (ResultSender, ResultReceiver) = mpsc::channel(100);\n```\n\n[src/asynchronous/client.rs @ f31f592, Client::new_stream](https://github.com/containerd/ttrpc-rust/blob/f31f5925749bba2616bc53942dd83877a3f5b532/src/asynchronous/client.rs#L148)\n\nWith the reader doing that send itself, one full channel stops the reader, and every other call on the connection waits behind it. Here is the reviewer's reproduction: a 200-frame stream that nobody reads, plus an unrelated unary call.\n\n*Interactive: [the 200-frame reproduction](https://notes.shvbsle.in/ttrpc-data-loss/#naive-fix).*\n\nEach call gets its own mailbox: an unbounded queue that only the reader writes to.\n\n*Interactive: [the same reproduction with mailboxes](https://notes.shvbsle.in/ttrpc-data-loss/#the-better-fix).*\n\nThe server also hands frames to tokio tasks, so why did only the client need this fix?\n\nThis is my PR that this note is based on: [containerd/ttrpc-rust#312](https://github.com/containerd/ttrpc-rust/pull/312)[↩](#fnref-1)\n\n\"GRPC for low-memory environments.\" The README explains that grpc-go's memory overhead is a problem \"when running a large number of services on a single machine\". From [containerd/ttrpc](https://github.com/containerd/ttrpc).[↩](#fnref-2)\n\n\"The protocol does not include features for handling unreliable connections such as handshakes, resets, pings, or flow control.\" From the [ttrpc protocol specification](https://github.com/containerd/ttrpc/blob/main/PROTOCOL.md).[↩](#fnref-3)\n\nJohn Salvatier, [Reality has a surprising amount of detail](http://johnsalvatier.org/blog/2017/reality-has-a-surprising-amount-of-detail), 2017.[↩](#fnref-4)\n\n\"The ordering bug is real, but this implementation introduces connection-wide head-of-line blocking. [...] I reproduced this by leaving a 200-frame stream unread and issuing an unrelated unary RPC: it times out with this patch.\" Tim Zhang, [review on #312](https://github.com/containerd/ttrpc-rust/pull/312#pullrequestreview-4863045230).[↩](#fnref-5)\n\nA stress test on the merged code saw about 1 in 2000 client-streaming calls with two adjacent frames swapped. None lost a frame.[↩](#fnref-6)", "url": "https://wpnews.pro/news/data-loss-in-ttrpc-rust", "canonical_source": "https://shvbsle.in/data-loss-in-ttrpc-rust/", "published_at": "2026-10-04 07:29:00+00:00", "updated_at": "2026-10-04 07:42:32.836567+00:00", "lang": "en", "topics": ["ai-agents", "developer-tools", "ai-infrastructure"], "entities": ["ttrpc-rust", "containerd", "tokio", "Rust", "Golang", "runc", "issue #311", "src/asynchronous/client.rs"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/data-loss-in-ttrpc-rust", "markdown": "https://wpnews.pro/news/data-loss-in-ttrpc-rust.md", "text": "https://wpnews.pro/news/data-loss-in-ttrpc-rust.txt", "jsonld": "https://wpnews.pro/news/data-loss-in-ttrpc-rust.jsonld"}}