{"slug": "can-you-use-autoregressive-diffusion-to-generate-market-data", "title": "Can you use autoregressive diffusion to generate market data?", "summary": "Jane Street research intern Kavish Kondap built an event-level generative model of US equities market data using autoregressive diffusion, trained on four years of US equities data with each row containing timestamp, price, kind (trade, order, cancel) and other information. Kondap found that DDPM diverged despite being theoretically ideal for denoising, while flow matching performed much better, and that fully continuous diffusion did not lend itself to the jaggedness of real market data. The attempts to handle point masses via smoothing and by discretizing some of the model's targets clarified what an eventual generative model of market data might look like.", "body_md": "*The following is part of a series of posts about 2026 summer intern projects – for more, see [“What the interns have wrought, special jumbo 2026 edition”](https://blog.janestreet.com/wrought-2026/)*\n\nIn quantitative finance we are used to models that take a stream of market data events for\na given symbol (like resting orders being added to an order book, cancellations of those\norders, executions) and predict that symbol’s future price. *Generative* models are less\ncommon. Imagine a model that could give not just a point estimate of a symbol’s price, but\ncould actually synthesize book events—including the timing of their arrival on the\nexchange. You’d get price predictions, of course, but your rollouts would have hugely more\ntexture than just that.\n\nBut what kind of data is market data? A generative model demands an answer to that\nquestion. Is it more like video, or text? That is, is it continuous or discrete? Market\ndata seems to have features of both. An order book evolves via a series of turns as market\nparticipants put on, or take away, resting orders; there is no doubt a discrete action\nspace. And yet, many of the most important parameters of a given order, like its price,\nhave such a high cardinality that they basically appear continuous. Some simply *are*\ncontinuous: an easy one to overlook, but crucial for a generative model, is timing, as in\n“when does the new order arrive at the exchange.” Even that oversimplifies things. In\npractice the distributions are spiky: “pennying” (where you improve upon a price by one\ntick) is much more common than improving by *two* ticks, and orders tend to arrive in\nbursts, for instance around whole-number times.\n\nOne way to explore these complications is to treat your data *as if* it’s continuous and\nsee how that breaks down. This summer, a research intern, Kavish Kondap, built a diffusion\nmodel for market data. Diffusion and flow-matching models are powerful tools for\nrepresenting sequential, multimodal data in regimes such as\n[image](https://arxiv.org/abs/2006.11239), [video](https://arxiv.org/abs/2204.03458),\n[robotics](https://arxiv.org/abs/2303.04137v5), and\n[audio](https://arxiv.org/abs/2309.07195). Taking inspiration from [*Autoregressive Image\nGeneration without Vector Quantization*](https://arxiv.org/abs/2406.11838), Kavish created\nan event-level generative model of market data using autoregressive diffusion.\n\nThere were some technical findings along the way—for example, Kavish found that DDPM, while theoretically ideal for denoising in diffusion models, diverged, and that flow matching performed much better. But the headline result was that fully continuous diffusion did not lend itself to the jaggedness of real market data. Kavish’s various attempts to handle point masses via smoothing, on the one hand, and discretizing some of the model’s targets, on the other, have greatly clarified what an eventual generative model of market data might look like.\n\n## Modeling and setup\n\nKavish looked at four years of US equities data, with each row of data containing timestamp, price, kind (trade, order, cancel, etc.), and other information. The diffusion model would try to generate these features for the next event, which was then autoregressively appended to the stream.\n\nFollowing the Li et al. paper cited above, he used an encoder–diffuser architecture,\nspecifically a causally masked transformer encoder to produce a latent embedding at each\nposition. During training, the latent at each position was fed into a small event-kind\nhead, which outputs a 2-categorical probability distribution indicating whether the next\nevent is a trade or a BBO update (“best bid and offer”). The continuous targets, like the\nprice, elapsed time, etc., are generated via diffusion head that is conditioned on the\ncorresponding latent and the *true* event-kind via standard [AdaLN\nconditioning](https://arxiv.org/abs/2212.09748).\n\nDuring inference, you pass only the encoded latent of the last event into the event-kind\nhead and sample an event from the output distribution. The next event’s continuous\nfeatures are generated via the diffusion head, conditioned on the last event’s latent and\nthe *sampled* event.\n\n## DDPM vs. Flow matching\n\nA DDPM or “denoising diffusion probabilistic model” predicts the noise that was added to a sample, and derives a clean estimate by subtracting the scaled noise prediction and dividing by the remaining signal level. DDPMs are highly successful in areas such as image generation, but at high noise levels, small errors in the noise prediction result in large errors in the clean target estimate. Flow matching instead interpolates linearly between noise and data, and trains the network to predict the velocity along that line (data minus noise). As a result, the sampled trajectories are nearly straight, and flow models can be integrated accurately in relatively few steps.\n\nFundamentally these are two different parameterizations of the same objective, but Kavish\nfound that the differences mattered quite a bit. His initial diffusion head predicted\n on a 1,000 step cosine schedule, as recommended in [Improved\nDDPM](https://arxiv.org/abs/2102.09672). At inference time, generation was run following\n[DDIM](https://arxiv.org/abs/2010.02502). Unfortunately, this configuration was unstable,\nand led to exploding denoising trajectories, with 88-95% of values over 8 standard\ndeviations away from the distribution mean on all 7 continuous targets. Note in the figure\nhow in the leftmost box, the trajectories (those thin pink filaments, just barely visible\nagainst the background) quickly diverge to the edges.\n\nThis is caused by the scaling of when calculating . As the figure shows, using more sampling steps reduces the increment per step, which prevents the amount of overflow. Additionally, adding stochasticity to the sampling process () regularizes the denoising process to a standard gaussian, reducing deviations. But while it is possible to tune the noise schedule to account for this instability, or apply inference-time interventions to prevent denoising explosion, Kavish found that a rectified flows approach performed well out-of-the-box. So he switched to a flow-matching model.\n\n## Handling discontinuities\n\nMost of the project was spent grappling with a fundamental problem: market data is neither fully continuous nor fully discrete. Video, for instance—for which diffusion models are well suited—is continuous both in snapshots (each frame is a series of continuous pixel intensities and colors) and in evolution (pixels can take on any continuous value from frame to frame). Market data, by contrast, has continuous snapshots but evolves in a very discrete way, governed by market microstructure.\n\nThis is evident when you look at the distributions of features, which are characterized by sharp boundaries:\n\nThe timing of orders isn’t normally distributed, but rather quite clumpy, with point masses around zero seconds (events that arrive simultaneously), the-earliest-possible-reaction-time, and whole-number seconds. And prices tend to cluster around the midpoint of the bid and ask, or the current bid, or the current ask, and thus you see sharp discontinuities there too.\n\nBecause there’s also an imbalance in categorical predictions—only 8% of events are trades, the rest being BBO changes; thus the categorical head overpredicts trades—Kavish experimented with fracturing the number of classes predicted by the categorical head. He broke BBO changes into a bunch of common classes, like “size only” (no change in price) or “ask up”, “bid down”, etc. This had the added benefit of allowing him to factor out “atoms” (or sharp discontinuities) in certain features, so that an elapsed time of zero could be predicted as its own class. He ended up with a 20-class categorical head, and simpler diffusion targets, like “how big was the non-zero time gap” (since zeroes are accounted for categorically) and “what’s the magnitude of the bid price change” (since direction is a category).\n\nThis approach led to materially better marginal distributions for both the event-kind and continuous targets, but obviously involved lots of hand-engineering. Using a categorical head to specially model every discontinuity in every feature becomes unscalable in larger feature sets. For instance, real data is likely to have far more variables, including high-cardinality ones like “which exchange was this on?”, potentially exploding the number of categorical classes.\n\nSo Kavish went looking for another way to handle discontinuities.\n\n## Atom smoothing\n\nHe used a procedure he ended up calling “atom smoothing.” The idea is to take the true, spiky distribution, smooth it out, and then re-sculpt in the spike, in a way that captures the probability mass there without creating a discontinuity. For example:\n\nFor this to work well, you have to be sure to pick the right representation of your variables: sharp spikes under some representation will be less sharp under another, and you need to expose and smooth any sharp spikes.\n\n## Results and rollouts\n\nKavish found that the diffusion model with flow matching and atom smoothing worked pretty\nwell. He compared the marginal distribution of the model’s output, per target, against the\ntrue marginal distribution. This was done in a one-step setting: he conditioned on a real\ncontext of 10,240 events, drew a single next event from the model, and compared it to the\nevent that actually occurred. He quantified the deviation from the true per-target\nmarginal distribution via total variational (TV) divergence, calculated as\n\n between two distributions\n and .\n\nAlthough alignment of the marginal distributions is important, it doesn’t capture joint relationships between features, or temporal consistency with prior context. So Kavish also trained a classifier to discriminate between model-generated predictions and real events. The classifier reads a window of real context events followed by a single candidate next event, and predicts whether that candidate was drawn from the model or from the data. The generator with feature representations optimized for atom smoothing solicited much more realistic samples:\n\nOf course the ultimate goal of the model was to produce autoregressive rollouts:\n\nThe model’s predictions naturally degrade the further into the future you get, and one question for further research is: by how much? In the second plot you can see how the spread in CME US widens in the synthetic rollout.\n\nKavish’s model is not yet accurate enough to be a realistic generator, but his research during the internship clarified what the important knobs are: how much discreteness should you bake into your model? (I.e., what exactly should you predict categorically, versus with diffusion?) And how do you represent and smooth the distribution of features to give diffusion its best chance?\n\nThe project helped us conceive of market data book events as simultaneously existing in a continuous space (“ask size increase of 10%”) and in a discrete action space (“penny the bid”). These are two ways of seeing the same thing. If you want to generate market data you have to model that blend correctly.\n\n## Looking forward to next summer…\n\n*If you’re interested in doing work like this, consider applying! You can find more details here: [Jane Street Internships](https://www.janestreet.com/join-jane-street/internships/). Applications for our 2027 ML Research internship are now open!*", "url": "https://wpnews.pro/news/can-you-use-autoregressive-diffusion-to-generate-market-data", "canonical_source": "https://blog.janestreet.com/can-you-use-autoregressive-diffusion-to-generate-market-data/", "published_at": "2026-09-30 00:00:00+00:00", "updated_at": "2026-09-30 23:18:31.306893+00:00", "lang": "en", "topics": ["machine-learning", "generative-ai", "ai-research"], "entities": ["Jane Street", "Kavish Kondap", "DDPM", "Li et al."], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/can-you-use-autoregressive-diffusion-to-generate-market-data", "markdown": "https://wpnews.pro/news/can-you-use-autoregressive-diffusion-to-generate-market-data.md", "text": "https://wpnews.pro/news/can-you-use-autoregressive-diffusion-to-generate-market-data.txt", "jsonld": "https://wpnews.pro/news/can-you-use-autoregressive-diffusion-to-generate-market-data.jsonld"}}