The DATA Foundation Launches to Tackle AI’s Multi-Billion Dollar Training Data Bottleneck
AI’s training data bottleneck has reached a critical juncture, and today, Story announces a strategic transition to become The DATA Foundation (“DATA”). The organization launches the DATA Network, anchored by a flagship integration with Kled, the world’s largest opt-in human data marketplace, registering 1.5 billion user-contributed records. Additionally, Trace is introduced as an onchain registry for AI training data provenance and licensing. Andrea Muttoni becomes CEO of The DATA Foundation, and Kled’s founder, Avi Patel, joins as Chief Data Officer in an advisory capacity.
AI training data has emerged as the most valuable and least solved category of intellectual property. Frontier AI labs have encountered a multi-billion-dollar data bottleneck, having effectively exhausted the internet for scraping. Remaining data is either expensive, bespoke, or legally undocumented, leaving labs unable to source data at scale, prove its provenance, or guarantee quality. Legal stakes are rising as labs rely on opaque networks without clear records of consent or jurisdiction. Scraped and undocumented data is no longer viable for enterprise-grade AI.
“The challenge in AI has shifted from compute and architecture to sourcing and provenance. As the scrapable web fractures, the question for labs now is who is keeping the receipts,” said Andrea Muttoni. “With Kled, we combine full data transparency and auditability with the largest pool of AI training data on the planet.”
DATA builds on Story’s original mission to deliver a data and IP layer for the internet, recognizing that the most critical form now is AI training data. The DATA Network brings essential infrastructure, starting with Kled’s licensing rails and contributor receipts, which now include stable coin payouts. This registration involves a staggering 1.5 billion user-contributed records with programmatic legal safeguards.
Avi Patel commented, “Frontier labs have exhausted the supply of high-quality, human-generated public text available on the open web. Suppliers showing data-sourcing provenance will win the next decade of deals, and that’s our bet. Instead of sourcing data blindly, Kled’s data marketplace and DATA’s auditable chain of custody converge on what labs actually need to license data with confidence and transparency.”
Trace, the public audit and search platform, generates immutable, confidential receipts for every contribution, allowing labs to verify dataset legitimacy in seconds. For every record uploaded, a receipt on DATA is generated, enabling upstream compensation for contributors’ data and intellectual property. This addresses the urgent need for a verifiable and compliant AI training data market.
DATA’s thesis was validated by Poseidon, the AI data processing project incubated by Story, which cleans, normalizes, and scores raw human data. Backed by a16z and running entirely on DATA, Poseidon’s contributor app Numo brings thousands of contributors into the AI economy with real-time payouts.
“We started Story to build an IP layer for the internet, and the most important IP of this era is the data you can’t scrape: how a surgeon’s hands move, how a robot grips, how people speak, drive, and work in the real world,” said SY Lee, CEO of PIP Labs and strategic adviser to The DATA Foundation. “DATA is where that conviction goes next: an end-to-end network that proves real-world data’s origin, licenses it, and pays the people who made it.”
The $IP token migrates to $DATA one-to-one with no action required from existing holders. Migration guidance and an FAQ are available at datafdn.org.