CentOS founder launches OpenWALDO, an open-source AI training dataset
Gregory Kurtzer, the founder of CentOS and Rocky Linux, launched OpenWALDO this week, a project building a shared, transparent AI training dataset that anyone can contribute to, the same collaborative model that built Linux. Kurtzer's company CIQ funds the project alongside Rocky Linux.
The pitch is provenance. Even open-weight models are typically trained on closed datasets pulled from copyrighted material, distilled outputs of other models, and scraped user content with no consent trail attached. OpenWALDO wants a dataset where every source is documented and inspectable, closer to how a Linux distribution's package sources are auditable end to end.
The dataset currently holds 167.3 billion reference tokens, drawn from government records, academic papers, mailing lists, and public domain literature. Kurtzer himself called that "a drop in the bucket next to the tens of trillions" frontier labs train on, and closing that gap is the open question the project doesn't yet answer.
For teams evaluating open-weight models against proprietary ones, an auditable training corpus does not by itself close the capability gap. What it offers is something proprietary models cannot: a way to check what a model was actually trained on, rather than taking a vendor's data-sourcing claims on faith. Whether that matters enough to draw contributors away from labs racing on raw scale is the real test for OpenWALDO over the next year.