Can you use autoregressive diffusion to generate market data? Jane Street research intern Kavish Kondap built an event-level generative model of US equities market data using autoregressive diffusion, trained on four years of US equities data with each row containing timestamp, price, kind (trade, order, cancel) and other information. Kondap found that DDPM diverged despite being theoretically ideal for denoising, while flow matching performed much better, and that fully continuous diffusion did not lend itself to the jaggedness of real market data. The attempts to handle point masses via smoothing and by discretizing some of the model's targets clarified what an eventual generative model of market data might look like. The following is part of a series of posts about 2026 summer intern projects – for more, see “What the interns have wrought, special jumbo 2026 edition” https://blog.janestreet.com/wrought-2026/ In quantitative finance we are used to models that take a stream of market data events for a given symbol like resting orders being added to an order book, cancellations of those orders, executions and predict that symbol’s future price. Generative models are less common. Imagine a model that could give not just a point estimate of a symbol’s price, but could actually synthesize book events—including the timing of their arrival on the exchange. You’d get price predictions, of course, but your rollouts would have hugely more texture than just that. But what kind of data is market data? A generative model demands an answer to that question. Is it more like video, or text? That is, is it continuous or discrete? Market data seems to have features of both. An order book evolves via a series of turns as market participants put on, or take away, resting orders; there is no doubt a discrete action space. And yet, many of the most important parameters of a given order, like its price, have such a high cardinality that they basically appear continuous. Some simply are continuous: an easy one to overlook, but crucial for a generative model, is timing, as in “when does the new order arrive at the exchange.” Even that oversimplifies things. In practice the distributions are spiky: “pennying” where you improve upon a price by one tick is much more common than improving by two ticks, and orders tend to arrive in bursts, for instance around whole-number times. One way to explore these complications is to treat your data as if it’s continuous and see how that breaks down. This summer, a research intern, Kavish Kondap, built a diffusion model for market data. Diffusion and flow-matching models are powerful tools for representing sequential, multimodal data in regimes such as image https://arxiv.org/abs/2006.11239 , video https://arxiv.org/abs/2204.03458 , robotics https://arxiv.org/abs/2303.04137v5 , and audio https://arxiv.org/abs/2309.07195 . Taking inspiration from Autoregressive Image Generation without Vector Quantization https://arxiv.org/abs/2406.11838 , Kavish created an event-level generative model of market data using autoregressive diffusion. There were some technical findings along the way—for example, Kavish found that DDPM, while theoretically ideal for denoising in diffusion models, diverged, and that flow matching performed much better. But the headline result was that fully continuous diffusion did not lend itself to the jaggedness of real market data. Kavish’s various attempts to handle point masses via smoothing, on the one hand, and discretizing some of the model’s targets, on the other, have greatly clarified what an eventual generative model of market data might look like. Modeling and setup Kavish looked at four years of US equities data, with each row of data containing timestamp, price, kind trade, order, cancel, etc. , and other information. The diffusion model would try to generate these features for the next event, which was then autoregressively appended to the stream. Following the Li et al. paper cited above, he used an encoder–diffuser architecture, specifically a causally masked transformer encoder to produce a latent embedding at each position. During training, the latent at each position was fed into a small event-kind head, which outputs a 2-categorical probability distribution indicating whether the next event is a trade or a BBO update “best bid and offer” . The continuous targets, like the price, elapsed time, etc., are generated via diffusion head that is conditioned on the corresponding latent and the true event-kind via standard AdaLN conditioning https://arxiv.org/abs/2212.09748 . During inference, you pass only the encoded latent of the last event into the event-kind head and sample an event from the output distribution. The next event’s continuous features are generated via the diffusion head, conditioned on the last event’s latent and the sampled event. DDPM vs. Flow matching A DDPM or “denoising diffusion probabilistic model” predicts the noise that was added to a sample, and derives a clean estimate by subtracting the scaled noise prediction and dividing by the remaining signal level. DDPMs are highly successful in areas such as image generation, but at high noise levels, small errors in the noise prediction result in large errors in the clean target estimate. Flow matching instead interpolates linearly between noise and data, and trains the network to predict the velocity along that line data minus noise . As a result, the sampled trajectories are nearly straight, and flow models can be integrated accurately in relatively few steps. Fundamentally these are two different parameterizations of the same objective, but Kavish found that the differences mattered quite a bit. His initial diffusion head predicted on a 1,000 step cosine schedule, as recommended in Improved DDPM https://arxiv.org/abs/2102.09672 . At inference time, generation was run following DDIM https://arxiv.org/abs/2010.02502 . Unfortunately, this configuration was unstable, and led to exploding denoising trajectories, with 88-95% of values over 8 standard deviations away from the distribution mean on all 7 continuous targets. Note in the figure how in the leftmost box, the trajectories those thin pink filaments, just barely visible against the background quickly diverge to the edges. This is caused by the scaling of when calculating . As the figure shows, using more sampling steps reduces the increment per step, which prevents the amount of overflow. Additionally, adding stochasticity to the sampling process regularizes the denoising process to a standard gaussian, reducing deviations. But while it is possible to tune the noise schedule to account for this instability, or apply inference-time interventions to prevent denoising explosion, Kavish found that a rectified flows approach performed well out-of-the-box. So he switched to a flow-matching model. Handling discontinuities Most of the project was spent grappling with a fundamental problem: market data is neither fully continuous nor fully discrete. Video, for instance—for which diffusion models are well suited—is continuous both in snapshots each frame is a series of continuous pixel intensities and colors and in evolution pixels can take on any continuous value from frame to frame . Market data, by contrast, has continuous snapshots but evolves in a very discrete way, governed by market microstructure. This is evident when you look at the distributions of features, which are characterized by sharp boundaries: The timing of orders isn’t normally distributed, but rather quite clumpy, with point masses around zero seconds events that arrive simultaneously , the-earliest-possible-reaction-time, and whole-number seconds. And prices tend to cluster around the midpoint of the bid and ask, or the current bid, or the current ask, and thus you see sharp discontinuities there too. Because there’s also an imbalance in categorical predictions—only 8% of events are trades, the rest being BBO changes; thus the categorical head overpredicts trades—Kavish experimented with fracturing the number of classes predicted by the categorical head. He broke BBO changes into a bunch of common classes, like “size only” no change in price or “ask up”, “bid down”, etc. This had the added benefit of allowing him to factor out “atoms” or sharp discontinuities in certain features, so that an elapsed time of zero could be predicted as its own class. He ended up with a 20-class categorical head, and simpler diffusion targets, like “how big was the non-zero time gap” since zeroes are accounted for categorically and “what’s the magnitude of the bid price change” since direction is a category . This approach led to materially better marginal distributions for both the event-kind and continuous targets, but obviously involved lots of hand-engineering. Using a categorical head to specially model every discontinuity in every feature becomes unscalable in larger feature sets. For instance, real data is likely to have far more variables, including high-cardinality ones like “which exchange was this on?”, potentially exploding the number of categorical classes. So Kavish went looking for another way to handle discontinuities. Atom smoothing He used a procedure he ended up calling “atom smoothing.” The idea is to take the true, spiky distribution, smooth it out, and then re-sculpt in the spike, in a way that captures the probability mass there without creating a discontinuity. For example: For this to work well, you have to be sure to pick the right representation of your variables: sharp spikes under some representation will be less sharp under another, and you need to expose and smooth any sharp spikes. Results and rollouts Kavish found that the diffusion model with flow matching and atom smoothing worked pretty well. He compared the marginal distribution of the model’s output, per target, against the true marginal distribution. This was done in a one-step setting: he conditioned on a real context of 10,240 events, drew a single next event from the model, and compared it to the event that actually occurred. He quantified the deviation from the true per-target marginal distribution via total variational TV divergence, calculated as between two distributions and . Although alignment of the marginal distributions is important, it doesn’t capture joint relationships between features, or temporal consistency with prior context. So Kavish also trained a classifier to discriminate between model-generated predictions and real events. The classifier reads a window of real context events followed by a single candidate next event, and predicts whether that candidate was drawn from the model or from the data. The generator with feature representations optimized for atom smoothing solicited much more realistic samples: Of course the ultimate goal of the model was to produce autoregressive rollouts: The model’s predictions naturally degrade the further into the future you get, and one question for further research is: by how much? In the second plot you can see how the spread in CME US widens in the synthetic rollout. Kavish’s model is not yet accurate enough to be a realistic generator, but his research during the internship clarified what the important knobs are: how much discreteness should you bake into your model? I.e., what exactly should you predict categorically, versus with diffusion? And how do you represent and smooth the distribution of features to give diffusion its best chance? The project helped us conceive of market data book events as simultaneously existing in a continuous space “ask size increase of 10%” and in a discrete action space “penny the bid” . These are two ways of seeing the same thing. If you want to generate market data you have to model that blend correctly. Looking forward to next summer… If you’re interested in doing work like this, consider applying You can find more details here: Jane Street Internships https://www.janestreet.com/join-jane-street/internships/ . Applications for our 2027 ML Research internship are now open