A common language for what AI is built from.
A shared account of the datasets and models the field runs on: where each came from, what is known about its terms, what it derives from, and what derives from it. We're building it with the institutions that have kept the web open, and publishing the design as we go, so the people it describes can shape it while that's still easy.
Four things about every entry.
- Origin: who published it, when, and where it came from.
- Terms: what the asset claims about its own permissions, in whatever license it already uses.
- Lineage: known derivations, upstream and downstream.
- Gaps: what we could not establish, stated plainly rather than left blank.
The Registry will record the terms each asset claims. It will not adjudicate between them. Roughly twenty licensing vocabularies now compete in this space with no agreement on which governs what. Recording what each asset says about itself does not require solving that first.
Four things it doesn't do.
- Not a standard: We are not asking anyone to adopt anything.
- Not for sale: Position in the record cannot be bought. Nobody pays to rank higher, appear sooner, or disappear.
- Not a gate: No entry confers permission and no absence withholds it.
- Not complete: It won't be. Gaps will be marked as gaps.
Compiled by us. Correctable by anyone.
No enrollment
We compile the record from what is already public. Nobody has to join it, sign anything, or opt in for it to be useful. An index is a claim about the world, not a membership list.
Every record challengeable
Every entry will carry a challenge link. Challenges will be recorded, attributed, and dated alongside the entry, whether or not we agree with them. Nothing gets quietly amended. This is the commitment that makes a compiled record legitimate rather than presumptuous, and it ships with the first version.
Yours to claim
If you publish or steward an asset in the index, you will be able to claim its record: confirm what's right, amend what isn't, add what we missed. Claimed records will be marked as claimed.
Not ours to own.
Archives, libraries, universities, Creative Commons, Wikimedia and the open source communities have spent decades keeping this material available. They already describe it, each in their own way. The work is to agree on a shared way of saying it, and then to make sure the result can be mirrored anywhere rather than held in one place.
AI Commons is the place that this conversation happens.
The Global Data Pledge.
Nobody has to sign anything for this record to exist. We compile it from what is already public. Commitment is a separate thing, and it already has a home.
The Global Data Pledge is a coalition of governments, memory institutions, publishers and rightsholder collectives committing to a consent-aware, provenance-rich commons. Its second ask (reciprocity) is recorded here, in this registry. That's the connection between the two: the Pledge is the commitment, and the registry is where the commitment becomes visible and stays visible.
Step three of four.
This initiative is currently underway with daily pledges, working groups drafting the technical standards and legal templates, and we plan to launch the live registry in Q1 2027. Model compliance will be ongoing work after that.