# Dev log #24 Hardening the p2p stack: From flaky tests to WebSocket fixes

> Source: <https://dev.to/yashksaini/dev-log-24-hardening-the-p2p-stack-from-flaky-tests-to-websocket-fixes-55f>
> Published: 2026-10-05 09:12:47+00:00

Hi, glad you found your way here. I'm Yash, a developer who contributes to open source, mostly p2p networking and AI tooling. Each week I write up what I shipped, what broke, and what fixing it taught me.

[GitHub](https://github.com/yashksaini-coder) · [X](https://x.com/0xcrackedDev) · [LinkedIn](https://www.linkedin.com/in/yashksaini) · [Portfolio](https://yashksaini.vercel.app/)

If you've ever had a CI run fail because a test expected a packet to arrive in exactly 200ms, you know my pain this week. I spent most of my time moving away from clock-based testing and toward progress-driven logic, alongside some critical bug fixes in the Python libp2p stack. Across 22 commits and 15 code reviews, it was a heavy week for p2p infrastructure and making sure our networking primitives actually behave when the network gets noisy.

Most of my energy went into **minip2p**, where I’m deep in the weeds of making the Rust implementation more resilient. I pushed 14 commits there, but the real story is in the refactoring.

There’s nothing that kills developer velocity quite like a flaky test suite. I opened and closed Issue #261 this week because our timing-budgeted tests were falling over whenever the CI runner got too busy. We’ve all been there: you put a `sleep` in a test, it works on your machine, and then it fails in GitHub Actions because the runner is under-provisioned.

I landed PR #284 to address this. Instead of waiting for the clock, I refactored the tests to be driven by progress. For the TCP transport tests, I stopped waiting for an arbitrary 200ms and started waiting for the socket to actually fill and drain. I also had to isolate the relay reservation and timer-bound tests in `nextest` to ensure they aren't fighting for resources. It’s the kind of unglamorous work that makes the whole project feel more professional.

Beyond the tests, I landed a pretty significant refactor in PR #260. I wanted to consolidate how we handle framed exchanges. Now, AutoNAT and Identify (two core p2p protocols) share a framed exchange and are properly bounded by their declared lengths. It’s a bit of "defensive networking"—ensuring a peer can't just stream infinite data at us during a handshake.

On the Python side of things, I spent some time in **py-libp2p**. I noticed a gnarly issue where the WebSocket transport was being a bit too generic with its error reporting. Basically, every failure after the initial connection was being swallowed and reported as a "handshake timeout," which is incredibly frustrating when you're trying to debug a real network issue. Even worse, it was leaking the TCP socket in the process.

I closed Issue #1554 by landing PR #1556, which ensures we report the *real* dial failure and, crucially, close the socket so we don't bleed file descriptors. I also followed up with PR #1555 to make sure we’re actually verifying server certificates on `wss://` (secure WebSocket) dials. Security isn't optional, even in p2p.

I also made a quick pit stop at **Trak**, a Rust project I've been helping out with. I noticed the indexer was getting distracted by test fixtures. If you have a directory full of mock code for tests, you don't necessarily want your primary code indexer treating those as real declarations. I landed PR #1 by updating the logic to read declarations from source code only and explicitly exclude test fixture trees. It’s a small change (+532/-12), but it makes the tool much more accurate for its users.

It was a "clean sweep" week for PRs—5 opened, 5 merged.

`sleep()` to actual state-checking.
While I only had 22 commits of my own, I spent a massive amount of time in the **dotnet-libp2p** repository. I provided 15 reviews this week, and honestly, this is where a lot of the "architectural" thinking happens. Helping the C# implementation stay in sync with the wider libp2p spec is a challenge, but the team there is shipping some great stuff.

Some of the highlights from the review pile:

`IDONTWANT` for large gossip messages to save bandwidth.
Reviewing 15 PRs while only opening 5 of my own definitely made this a "mentorship and alignment" week. It's important to keep the ecosystem moving together, especially when you're working across three different languages.

This week was a polyglot special:

`minip2p` and `Trak`.` py-libp2p` maintenance and bug fixing.
The numbers tell a story of growth: **+2,503 / -357**. That’s a lot of new code, but a good chunk of that was the test infrastructure and the refactored protocol handlers. I'm okay with a net positive line count when it means the tests are finally stable.

Next week, I want to double down on the `minip2p` transport layer. Now that the tests are stable, I can actually start benchmarking performance without worrying about the CI runner's mood. I'm also keeping an eye on the open PRs in the Python and .NET stacks—there's a lot of momentum right now in the p2p space, and I want to make sure we keep that energy going.

If you're building p2p tools, stop using `sleep()` in your tests. Your future self (and your CI) will thank you. See you next week!

**Generated by [DevNotion](https://github.com/yashksaini-coder/DevNotion)**
