Aurora Odds: should you go outside and look up right now? A developer built Aurora Odds, a browser-based tool that uses the open-weights TabPFN tabular foundation model to forecast whether the aurora will be visible from a user's exact location in the next hour. The model reads 2,000 labelled half-hours of solar wind data from 1998 to 2019 in context, using an arrival-time axis so each sample is shifted by its own travel time from the L1 point, and returns a full probability distribution for the coming hour's geomagnetic activity. The project includes a red night-vision mode, an alert that chimes when viewing odds pass 40%, and replay modes for the May 2024 and October 2024 storms in London and Chicago. This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass https://dev.to/challenges/hacktoberfest-week1-2026-10-05 On 10 May 2024 the northern lights reached Florida, northern India and the Canary Islands. Millions of people saw them. Many more only read about it the next morning, because the alerts on their phones said "G5" and "Kp 9" and nothing about their own sky. That is the gap I wanted to close. Every aurora app I tried answers a forecaster's question how disturbed is the magnetic field? when the question people actually ask is a practical one: is it worth putting my shoes on and going outside, here, right now? Aurora Odds https://luoy16002-svg.github.io/aurora-odds/ answers exactly that. Pick your place or let the browser find it and you get one verdict for your sky, Go outside now , Bring your phone , Not yet: it's still light out , or Stay in, the sky is quiet , backed by four numbers: There is also a red night-vision mode, because the whole point is to use it outside in the dark without wrecking your eyes' adaptation, and a "tell me when to go out" switch: leave the tab open and it chimes once it is dark, mostly clear and the odds for your eyes pass 40%. The screen can stay in your pocket until then. The forecast comes from TabPFN https://github.com/PriorLabs/TabPFN , an open-weights tabular foundation model, reading the solar wind that is on its way to Earth. Live: https://luoy16002-svg.github.io/aurora-odds/ https://luoy16002-svg.github.io/aurora-odds/ The page is most fun during a storm, and storms are rare, so it has a replay mode that runs the same model on the solar wind of a real night: London, 10 May 2024 https://luoy16002-svg.github.io/aurora-odds/?replay=may2024&place=London and Chicago, 10-11 October 2024 https://luoy16002-svg.github.io/aurora-odds/?replay=oct2024&place=Chicago . Clouds in a replay come from Open-Meteo's historical archive for that hour. Should you go outside and look up right now? Aurora Odds answers that for your exact sky: the chance the aurora is bright enough for your eyes or your phone camera in the next hour, whether it is dark and clear enough and which way to face and how high to look. Live page: https://luoy16002-svg.github.io/aurora-odds/ https://luoy16002-svg.github.io/aurora-odds/ · storm replays: London, 10 May 2024 https://luoy16002-svg.github.io/aurora-odds/?replay=may2024&place=London , Chicago, 10-11 October 2024 https://luoy16002-svg.github.io/aurora-odds/?replay=oct2024&place=Chicago The forecast is made by TabPFN https://github.com/PriorLabs/TabPFN , an open-weights tabular foundation model reading the solar wind measured at the L1 point 1.5 million km upstream. It is not trained on this problem: it reads 2,000 labelled half-hours from 1998 to 2019 in context 8,000 in the evaluation below and returns a full probability distribution for the coming hour's geomagnetic activity. One distribution answers every latitude. Spacecraft parked at the L1 point, 1.5 million kilometres sunward, measure the solar wind every minute. That wind takes 20 to 80 minutes to reach Earth, so part of the next hour is already measured before it arrives. The most important number is Bz , the north-south direction of the wind's magnetic field. When it points south it connects to Earth's field and pours energy in; when it points north, very little gets through, however fast the wind blows. I wanted the model to see the world exactly as it can be seen at one moment, so every input is built on an arrival-time axis : each sample is shifted by its own travel time, and at time t the model is only allowed to use samples that had already been measured at L1 by t . That gives four windows: what is measured but still on its way , the last 30 minutes, 30 to 90 minutes ago, and 90 to 180 minutes ago. In each I summarise Bz, the Newell coupling function https://doi.org/10.1029/2006JA012015 a standard estimate of how much energy the wind is delivering , speed, density and pressure, plus the season and time of day. The same code builds 28 years of history from NASA's OMNI data and the live features from NOAA's real-time feed, so there is no training/serving skew to debug at 2 a.m. during a storm. The target is the highest Hp30 in the coming hour. Hp30 is GFZ Potsdam's half-hourly version of the Kp index: same scale, finer timing, and open-ended, so a superstorm is not squashed at 9. TabPFN is a transformer pretrained on millions of synthetic tabular problems. You don't train it on your data: you hand it a table of labelled examples the context and it predicts new rows in a single forward pass, like in-context learning in a language model. I gave it 8,000 half-hours from 1998 to 2019 as context the live page uses 2,000; more on that below . Storms are rare the next hour reaches Hp30 6 only about 1.2% of the time , so a uniform sample shows the model very few of the moments that matter. I tried a stratified context that over-represents storms and then corrects the predicted distribution back to the real storm frequency. On the 2020-2022 validation years the stratified context won mean Brier skill 0.46 against 0.44 for Hp30 ≥ 5 to 7 , so that's what runs. The part I like most: TabPFN's regressor returns a full probability distribution over the next hour's Hp30, not a single number. So one forward pass answers every latitude at once. Tromsø needs almost nothing, Edinburgh needs about Kp 5 to see it with the naked eye, London about 7.5. Each of those is just P Hp30 ≥ k read from the same distribution. I tested on 2023 to 2026 , years the model never saw, which include the solar maximum, the May 2024 superstorm and the October 2024 storm. I compared it with: | Brier skill score AUC | Hp30 ≥ 4 | Hp30 ≥ 5 | Hp30 ≥ 6 | Hp30 ≥ 7 | Hp30 ≥ 8 | |---|---|---|---|---|---| | TabPFN, 8,000-row context | 0.56 0.951 | 0.54 0.975 | 0.52 0.987 | 0.58 0.993 | 0.53 0.990 | | TabPFN, 2,000-row context the live page | 0.54 0.948 | 0.52 0.973 | 0.51 0.986 | 0.56 0.992 | 0.50 0.992 | | LightGBM, trained on 370,431 rows | 0.56 0.951 | 0.54 0.974 | 0.51 0.987 | 0.46 0.989 | 0.34 0.982 | | Last half-hour's index oracle | 0.34 0.790 | 0.33 0.777 | 0.34 0.776 | 0.45 0.811 | 0.44 0.804 | 61,420 half-hours from January 2023 to September 2026. Storm half-hours at each level: 9,345, 3,458, 1,201, 485, 166. Brier skill is relative to always forecasting the long-run frequency; 0 means no skill. For moderate activity, TabPFN and LightGBM are level. For the rare, strong storms, the ones that bring the aurora to Edinburgh, Chicago or London, TabPFN is clearly better: a Brier skill of 0.58 against 0.46 at Hp30 ≥ 7 and 0.53 against 0.34 at Hp30 ≥ 8. LightGBM saw 46 times more rows, but strong storms are a sliver of them and a tree has very little to split on out there. My guess is that TabPFN's prior, learned from millions of synthetic problems, degrades more gracefully in the tail; either way, it's the tail that decides whether someone in London goes outside. The live page runs the 2,000-row context, so a free 4-core GitHub runner finishes in about a minute. That costs about 0.02 of skill. It is not magic. On 10 May 2024 the index climbed to about 6 in the early afternoon while the wind measured upstream still looked calm a weak field of about 3 nT, barely southward . A model that only sees the wind had nothing to react to and sat at Kp 2 to 3 until the storm's shock arrived at 17:10 UTC; after that it was in the right place within half an hour. It also over-forecasts a little at moderate levels: in the reliability chart the curves for Hp30 ≥ 5 and 6 sit slightly below the diagonal. For a tool whose only job is to tell you when to put your shoes on, I'm fine with erring that way. The rest is geometry, and it's where most aurora apps stop short. A GitHub Actions job reruns the model every ten minutes on a CPU runner and publishes the result with the static page. Three things in this project only work because the pieces are open. The model can be hammered for free, at exactly the wrong moment. Aurora traffic is the spikiest traffic there is: nobody looks for months, then a storm hits and everyone looks within the same ten minutes. Because TabPFN's weights are open, the model runs once every ten minutes in a free CI job and writes a small JSON file that a CDN serves to everyone. A million visitors cost the same as one. There is no API key in a secrets store, no per-call bill and no rate limit that kicks in during the one night that matters. I could change what the model returns, not just what I send it. The label-shift correction works on TabPFN's full predictive histogram: I re-weight its bars by how much more often storms appear in the context than in reality. That needs the raw distribution. An endpoint that returns one number would have made the stratified context unusable, and the "one distribution, every latitude" trick impossible. Every claim in this post can be checked. The solar wind history is NASA's OMNI data, the index is GFZ's Hp30 CC BY 4.0 , the live feed is NOAA's and the clouds are Open-Meteo's. With open weights on top, python pipeline/evaluate.py reproduces the numbers above on a laptop or a single consumer GPU. If you think my context is badly chosen, or that a different set of features would do better, you can try it in an afternoon and show me. Best Use of TabPFN : TabPFN is the forecaster. It reads the labelled history in context and its full predictive distribution turns into the chance of seeing the aurora at every latitude.