OpenWALDO Launches Open Training Data Corpus OpenWALDO, a Ctrl IQ Inc.-sponsored open-source project led by Gregory M. Kurtzer, launched publicly on July 29 with an August 11 index listing 167.3 billion reference tokens across 93.9 million documents, aiming to make AI training data reviewable and reproducible. The project proposes an AI Bill of Materials to link training inputs, licenses, and model outputs, though it has not released benchmark results or a foundation model. OpenWALDO Launches Open Training Data Corpus OpenWALDO went public on July 29 as a Ctrl IQ-sponsored project for a community-governed corpus of AI training data; SiliconANGLE covered the launch on August 11. The project publishes an AI Bill of Materials approach linking training inputs, licenses, artifacts and model outputs. Its August 11 public index listed 167.3 billion reference tokens across 93.9 million documents. OpenWALDO went public on July 29 as an open-source AI project sponsored by Ctrl IQ Inc. and led by Gregory M. Kurtzer. SiliconANGLE covered the launch on August 11. The project is building a community-governed corpus of AI training data and associated tooling intended to make model inputs reviewable, attributable and reproducible. OpenWALDO's public index listed 167.3 billion reference tokens across 93.9 million documents , divided into 46 corpora and 799 shards in its August 11 status snapshot. The index listed 51 asserted license identifiers. The project describes the corpus as a public record managed through Git review for metadata, content-addressed storage for canonical Parquet objects, and Developer Certificate of Origin, or DCO, sign-off for contributor responsibility. From open weights to traceable inputs The project frames its work around a distinction between open-weight models and fully open-source AI. OpenWALDO argues that downloadable model weights alone do not expose the data, methods and training recipes required to regenerate a model. SiliconANGLE reports that OpenWALDO proposes an AI Bill of Materials , or AI BOM, that records object references, documents, tokens, inventory and licenses associated with a canonical model. OpenWALDO describes its objective as carrying provenance from source data into downstream model artifacts. Kurtzer previously founded CentOS and Rocky Linux and created Singularity, now Apptainer, according to SiliconANGLE and his OpenWALDO introduction. Tooling and governance details OpenWALDO states that its core data and model workflows already run end to end in public, while contributors refine interfaces and add capabilities. The available project materials do not provide benchmark results, identify a released foundation model, or specify model-training partners. For ML teams, a versioned dataset inventory and documented license assertions can make data lineage more inspectable than a conventional opaque web-scale training mix. However, asserted license metadata is not equivalent to a legal determination of rights, data quality, or downstream model risk. Organizations evaluating any public corpus would still need to validate provenance, license terms, content policy, removal processes and compatibility with their own model-development and deployment requirements. OpenWALDO's approach places training data governance alongside source-code practices such as review, version control and attributable contributions. Comparable open-data efforts have often faced difficult tradeoffs between corpus scale, licensing certainty, deduplication, privacy review and the operational burden of processing corrections. The practical value of OpenWALDO's model will depend on whether its public records can support those requirements at useful scale for model builders. Key Points - 1OpenWALDO went public on July 29 and its August 11 index snapshot listed 167.3 billion reference tokens, expanding infrastructure for traceable open-source AI development. - 2Its AI Bill of Materials concept links data, licenses and artifacts, giving practitioners a proposed framework for training-data lineage review. - 3Comparable public corpora face persistent licensing, privacy, quality and correction challenges, so metadata transparency does not itself establish legal clearance. Scoring Rationale OpenWALDO introduces a substantial public training-data index and a concrete provenance-oriented governance model, both relevant to teams building reproducible open models. Its practitioner impact remains unproven because the project has not yet published model results or broad adoption evidence. Sources Primary source and supporting public references used for this report. Practice interview problems based on real data 1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with. Try 250 free problems /problems