{"slug": "what-the-interns-have-wrought-special-jumbo-2026-edition", "title": "What the interns have wrought, special jumbo 2026 edition", "summary": "Jane Street published a review of its 2026 summer intern projects, including ML research on an event-level generative model of market data built with autoregressive diffusion by intern Kavish Kondap and work by Monte Bohde on quantifying and mitigating LLM memorization using divergence decoding, distillation, natural language autoencoders, and prompt engineering. Other intern projects included Kohki Horie's neural network for per-order predictions across a variable-sized book, Alexander Gu's sequence model predicting stock returns and future input features, Ritvik Bale's KV cache distillation experiments to speed up prefill on a large model, and Dashiell Bhattacharyya's characterization of bf16 versus fp32 hidden-state degradation in a Mamba model. Jane Street said applications for its 2027 internship are now open.", "body_md": "It’s time again to review some of what Jane Street’s interns built over the summer, though this year we’re going to try to cover more projects and in more depth. In particular we’re including projects from teams and roles whose work has historically been hard to talk about publicly—like ML research and trading desk operations. There’s just been too much exciting work in these areas to leave them out.\n\nHopefully the below gives a more complete picture of the kind of stuff interns do here, with examples from ML Research, Linux Engineering, IT, Strategy & Product, and Trading Desk Operations Engineers. Even so it’s still a small sample out of hundreds of projects (and doesn’t talk at all about our trading internship, which is more of a you-had-to-be-there apprenticeship, and even harder to sum up).\n\n*If you’re interested in doing work like what you see below, consider applying! You can\nfind more details here: [Jane Street\nInternships](https://www.janestreet.com/join-jane-street/internships/). Applications for\nour 2027 internship are now open!*\n\nAnyway, let’s dive in.\n\n## ML Research\n\nIntern projects in ML research have a different flavor than our typical software\nengineering projects, in that they’re far more exploratory—they are, genuinely,\n*research*. Reviewing several interns’ work logs in preparing this post, I was struck by\njust how doggedly—and how quickly!—they perform experiments, often scaling up to\nserious training runs in the first week. (I’d have just gotten my editor set up the way I\nlike…) After a month and a half (interns do two projects each summer) they’re reporting\nfindings that, per the mentor of one of the projects below, can end up being “very\nimportant.”\n\n- \nKavish Kondap created an event-level generative model of market data using autogressive diffusion. If you could do long rollouts that fooled a classifier, you’d have a source of synthetic data. But even just the process of training such a model sheds light on a question that’s long interested us: just how continuous is market data? [Read the full writeup here](https://blog.janestreet.com/can-you-use-autoregressive-diffusion-to-generate-market-data/) .\n- \nMonte Bohde explored a variety of methods for combating memorization in LLMs. We want LLMs to make predictions, and in historical benchmarks they often seem quite good at it —but are they good at it because they’ve memorized the outcomes during pretraining? This is one reason such systems often fail on new prediction tasks in production, past the knowledge cutoff. Using divergence decoding, distillation, natural language autoencoders, and some prompt engineering, Monte sought to quantify memorization and see if he could shake it out of a model without hurting out-of-sample performance. [Read the full writeup here](https://blog.janestreet.com/mitigating-memorization-in-llms/) .\n\nThere were of course so many other cool projects that we couldn’t cover in depth, among them:\n\n- Kohki Horie’s neural network that can answer detailed questions about the future of any order on the book. Compared to a classical time-series setup, the interface that was being demanded—per-order predictions across a variable-sized book, all conditioned on a single shared input sequence—was unusual, and after much experimentation Khoki arrived at a clever architecture well-suited to it.\n- Alexander Gu’s sequence model that predicts both stock returns and future input features, then tests whether tree-search rollouts and self-distillation beat direct return prediction.\n- Ritvik Bale’s experiments to speed up prefill on a large model by warming it up with KV cache distillation.\n- Dashiell Bhattacharyya’s attempts to characterize bf16 vs. fp32 hidden-state degradation in a Mamba model to see if he could develop quantization-aware training.\n\nAnd others besides!\n\n## Software Engineering\n\nThere were a huge number of projects that we could have written up this year, but given that humans are doing the writing we had to keep it to a manageable number.\n\n- \nArsh Koneru developed an activation checkpointing scheme that outperformed PyTorch’s in order to keep model training runs within a tight RAM budget. In activation checkpointing, you recompute some nodes in a neural network’s backward pass instead of saving them from the forward pass. This trades off memory for compute. Arsh’s project involved finding a better set of nodes to recompute. He accomplished this by parsing the intermediate dynamo graph from `torch.compile` and greedily picking nodes with the\ncheapest recompute cost per byte of peak memory freed.[Read the full writeup here](https://blog.janestreet.com/trading-off-compute-for-memory-with-activation-checkpointing/) .\n- \nTheodor Totev built a performance testing framework and implemented an indexing strategy to speed up “tip recovery” in Aria, a messaging framework used heavily at Jane Street for [state machine\nreplication](https://signalsandthreads.com/state-machine-replication-and-why-you-should-care/) .\nWhen a client is slightly too slow to process messages, the server will still maintain\nthe connection as long as the client stays within one ring buffer capacity from the\n“tip” of the stream. Previously this meant scanning and filtering through all messages\nin the ring buffer even if the client was only interested in a much smaller subset of\nthe total Aria stream—this caused slowness in our servers on several occasions in\nproduction!  Theodor’s code builds up and maintains a dense index from each Aria “topic”\npartition into message offsets in the ring buffer, resulting in 30% less CPU utilization\nin real-world scenarios and significantly cutting overall latency across clients during\nthose times.\n- \nSai Konkimalla developed lib/spmd, an implicit single-program-multiple-data DSL in OxCaml targeting CPU SIMD instructions, similar to Intel’s ISPC. It takes inspiration from Halide and Dex by separating schedule from code and reasoning about index sets. Although Sai didn’t finish the SIMD backend, initial results against hand-written C intrinsics look promising.\n\nAnd just to give a quick shout to a few of the many other projects we would have loved to write more about:\n\n- \nMikołaj Kołek wrote an implementation of interning-/hash-consing-based compression of debug graphs in Datafetcher, a very powerful internal library for abstracting out requests for external data. Kołek’s project involved dealing with tricky races between multiple compressors working on the same graph.\n- \nStanislaw Malinowski evaluated ILP/SAT solvers as an alternative to our greedy scheduler for assigning market data applications to hosts, and got promising results.\n- \nChien-Yu Xiong implemented a better coordination protocol for controlling access to our Gurobi cluster, for linear programming and optimization problems.\n- \nOscar Xu improved the performance of an important feature engineering library that feeds into our trading systems. Previously, the system had to consume market data, keep track of features, and evaluate feature vectors all in one process. Oscar’s work allows some of this to happen in symbiote processes, exposing accessors to trading systems via symmetrically double-buffered addrs in shared memory.\n\n## Linux Engineering and IT\n\nLinux engineers at Jane Street have a broad remit, working on everything from datacenter provisioning and orchestration, kernels and drivers delivered through Nix and RPMs, to production and research storage at scale, managing cloud infrastructure that integrates cleanly with our on-prem platform, and, in partnership with the IT and Network Engineering teams, managing Linux desktops on the network. This summer, interns took on a slew of projects in these areas but we wanted to highlight a few that stood out:\n\n- \nJacob Root built a resilient kernel log reporter, ensuring that kernel logs would be available even during disk or network failures. To do this he had to root out surprising dependencies: even figuring out the host and port to send logs to relies on disk, because DNS resolution touches `/etc/resolv.conf` and`/etc/hosts` . When even the\nnetwork is down, Jacob’s tool falls back on Linux’s pstore subsystem, which provides a\nbare-bones persistent storage mechanism backed by the same non-volatile memory that\nstores system firmware. All this has allowed us to investigate individual crashes in the\nkind of depth we like, especially on systems without lights-out management that have\nhistorically been difficult to debug.\n- \nKian Kasad developed an OCaml-native FUSE (filesystem in userspace) library to build a high-performance read-only FUSE server for an in-house storage system called Depot. Historically, we’ve been able to rely on off-the-shelf storage products (like Dell Isilons or VAST Storage), but as dataset sizes have grown, we built Depot as a highly scalable object store specifically for research workloads. Unfortunately, various components in our core trading research infrastructure expect to access files via a regular POSIX filesystem interface, whereas Depot’s client library is all in OCaml. Kian’s project bridged the gap, and after some patient performance optimization reached read speeds of 1.3 GiB/s, with room still to improve.\n- \nMax Ohm developed a centralized configuration management system for Windows, replacing a suite of distributed PowerShell scripts that had become hard to reason about. In Max’s system, a central orchestrator, developed in OCaml, interacts with a stateless F# agent running on each host. The setup greatly simplifies the client-side logic and helps bring type safety to the top-level config scripts specifying which host classes to upgrade and how. One of the neat things about Max’s projects was the careful way it was tested, where the old and new systems could operate in tandem, with reports of any diffs between the two. This revealed some subtle bugs that had long lain dormant in the old system.\n\n## Trading Desk Operations Engineers\n\nTrading Desk Operations Engineers (TDOEs) keep our trading desks running smoothly. They collaborate deeply with traders, desk developers, and operations teams like finance and accounting. On a given day a TDOE might investigate an outage in a critical system, write OCaml code to automate an error-prone process, or sit with a trader to build a dashboard together. TDOEs must balance steadily making workflows and systems more resilient with the urgent needs of trading.\n\nThis summer’s internship had more projects across more teams than ever before. Interns’ work on the desks ranged from convertible-bond trade analysis with Global Capital Markets to cash flow reconciliation for exotic options trading.\n\nRead more—with more intern projects!—on the accompanying post: [What the TDOE interns have wrought, 2026 edition](https://blog.janestreet.com/tdoe-intern-projects-2026/)\n\n## Strategy and Product\n\nStrategy and Product Specialists are versatile, cross-functional contributors who drive forward firmwide initiatives. The role requires a unique combination of big-picture thinking (“why is this valuable to the business?”) and deep analytical problem-solving.\n\nFull-time SPs work across all areas of the firm: on trading desks, in infrastructure groups, and with tech teams. SP interns are embedded directly onto a team, paired with a full-time mentor, and given real-world projects we actually care about getting done. Some of this summer’s projects included automating fleetwide Linux kernel upgrades; modeling the fee expense on several types of loans to better allocate financing costs to different kinds of trading; and improving how we manage risk limits on unusual dates; and more.\n\nYou can read about them here: [What the SP interns have wrought](https://blog.janestreet.com/wrought-2026-sp/).", "url": "https://wpnews.pro/news/what-the-interns-have-wrought-special-jumbo-2026-edition", "canonical_source": "https://blog.janestreet.com/wrought-2026/", "published_at": "2026-09-30 00:00:00+00:00", "updated_at": "2026-09-30 23:18:21.198457+00:00", "lang": "en", "topics": ["machine-learning", "large-language-models", "ai-research", "artificial-intelligence"], "entities": ["Jane Street", "Kavish Kondap", "Monte Bohde", "Kohki Horie", "Alexander Gu", "Ritvik Bale", "Dashiell Bhattacharyya"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/what-the-interns-have-wrought-special-jumbo-2026-edition", "markdown": "https://wpnews.pro/news/what-the-interns-have-wrought-special-jumbo-2026-edition.md", "text": "https://wpnews.pro/news/what-the-interns-have-wrought-special-jumbo-2026-edition.txt", "jsonld": "https://wpnews.pro/news/what-the-interns-have-wrought-special-jumbo-2026-edition.jsonld"}}