I build iPulse AI at Future Edge Group. We have published a small, versioned dataset of historical consensus snapshots for people studying how to make AI investment research inspectable.
The Hugging Face viewer currently contains 746 rows. Each record keeps the snapshot identity, asset, forecast horizon, scoring timestamp, model/advisor counts, methodology versions and checksum together. The dataset is CC BY 4.0:
The distinction that matters: these are stored system outputs, not a validated trading benchmark. A consensus score is not a calibrated probability of being right, and agreement between advisors is not independent evidence of accuracy.
For reuse, I would keep snapshot IDs and version fields in the grouping keys, separate generation time from the forecast anchor, and avoid pooling horizons or algorithm versions into one performance number. Any outcome study still needs a declared evaluation window, realized outcomes, a baseline, and a treatment of missing or censored observations. The public product methodology is at iPulse AI Methodology Docs | AI Market Intelligence Transparency.
What metadata would you need before using a dataset like this for a leakage-safe evaluation? In particular, would you separate recorded outputs and realized outcomes into different tables, or require an immutable joined evaluation release?