Nothing records what AI is built from.
Every AI system that can write, reason, translate, or create learned to do so from the recorded work of human beings. There is no single place to see what that record contains, where it came from, or under what terms it was used. We're building one.
AI compressed humanity's memory and threw away the index.
Intelligence requires memory, and memory requires a record of where it came from. Civilization solved this a long time ago and solved it institutionally, with libraries, archives, citation, the patent record, version control. Each one governs what enters, what is kept, and how errors get corrected.
AI systems were trained on everything those institutions held and connected to none of them. The model keeps the content and discards the provenance. What it knows, it cannot trace. What it gets wrong, no one can follow back to the source.
Over 70% of widely used AI training datasets omit licensing information entirely.
What we don't count, we can't recognize.
Knowledge is cumulative. Every advance stands on contributions that came before it, most of them given freely and most of them never recorded. A modern smartphone rests on thousands of patents, a substantial share of them originating in publicly funded research labs. The last party in the chain is the only one anyone can see.
That is how enclosure actually happens. Not by theft, but by omission. When the record of contribution doesn't exist, the value of contribution accrues to whoever shipped last, and everyone upstream disappears from the story.
Open is not the same as neutral.
Openness is a property of artifacts. Neutrality is a property of the layer that organizes them.
You can keep every weight downloadable and still lose neutrality at the layer that decides what surfaces, what ranks, what gets recommended, what runs on whose compute, and what it costs. This is not a hypothetical failure. It is the pattern search ran, the pattern social platforms ran, and the pattern now arriving at the center of the AI ecosystem.
The open web survived because its foundations were not just open but neutral. Anyone could build on them, and nobody could quietly tilt them. Those are two different properties, and only the first one is currently being defended.
A layer is neutral when no participant can change its behavior in their own favor.
And when any change to how it behaves is visible to everyone who depends on it. In practice that means five things:
- Non-preference: Same input, same treatment, regardless of who is asking.
- Non-extraction: The operator does not profit from how the layer is arranged. If position can be sold, the layer is not neutral.
- Legibility: The rules are published and changes are logged. Neutrality you cannot inspect is a promise, not a property.
- Exit: Anyone can take their contribution and leave, or run their own copy.
- Non-capture: No participant or coalition can acquire control of it.
Neutrality is not a state you achieve. It is a condition you maintain, and it decays in a predictable direction: toward whoever gains most from a small exception. Preserving it is structural, not ethical. Publish the rules. Log the changes. Make sure the party with the most to gain from an exception is never the party who can authorize one.
A common language for where it came from.
The institutions that kept the web open (archives, libraries, universities, Creative Commons, Wikimedia, the open source communities) already hold this material and already describe it, each in their own way. What's missing is a shared way of saying it, so that a claim made in one place still means something in another.
We're building that with them. A common language, and a common record of what AI is built from. Held in the open rather than owned. Mirrored rather than centralized, because a record that lives in one place usually only answers to whoever owns that place.
Still legible and reliable, 10 years later.
Ten years from now, nobody will be asking whether AI was built from the world's knowledge. The question will be whether anyone can still see how.
A common language makes that visible. A record held in common keeps it visible after the funding cycles, the acquisitions and the policy moments have all moved on. That's the whole of it: knowledge that stays legible to the people it came from, and a coordinating layer that treats everyone running on it the same.
The Global Data Pledge is where that commitment gets made, by governments, memory institutions, publishers and rightsholder collectives, each at the level their holdings allow.
We've been convening this conversation for a decade.
AI Commons has worked across the AI for Good Summits, the G7 Hiroshima Process, the Global Partnership on AI, and a long line of convenings on what it means for AI to serve the people it came from. The index is the first thing we are putting on the record ourselves.