The Data Assetization Race: A Global Observation

The Data Assetization Race: A Global Observation

221 zettabytes. That’s roughly how much data the world is on track to generate in 2026 alone, according to Statista, up from about 181 zettabytes just the year before. It’s a number large enough that it stops meaning anything the moment you try to picture it. Every text message, every sensor reading, every video frame, multiplied a thousandfold and stacked on top of itself, year after year.

Scale was never the hard problem. The hard problem is what happens next: how much of that data ever gets an owner, a price, a way to keep paying the person who created it.

That gap, between data simply existing and data functioning as an asset, is exactly what the “data assetization” race is trying to close. Turn a byte from something a platform quietly holds into something with clear ownership, the ability to move, and the ability to generate ongoing income for whoever made it. The category has picked up real momentum over the past year or two, and it isn’t happening by accident. Three forces are pushing at once.

Driver One: AI’s Data Hunger

Bigger models need more data, that part barely needs explaining anymore. What’s worth watching is how fast that hunger is turning into a real market. Grand View Research puts the global AI training dataset market at roughly $3.2 billion in 2025, growing to $16.3 billion by 2033, a compound annual growth rate of 22.6%.

The demand side is shifting in an even more interesting direction. McKinsey estimates that by 2030, commerce initiated and executed autonomously by AI agents could reach $3 to $5 trillion globally, with the US market alone contributing as much as $1 trillion. Once the parties initiating, executing, and even negotiating a transaction are machines with no inherent basis for trust, the rules governing who owns a piece of data, whether it’s genuine, and how the resulting value gets split stop being a nice-to-have. They become the foundation the entire new economy has to run on.

Driver Two: Losing Control Is Now a Line Item

Not knowing who owns what, or being unable to control where data flows, used to read as an abstract compliance risk. Over the past couple of years it has turned into an actual bill.

IBM’s 2025 Cost of a Data Breach Report puts the global average cost of a single breach at $4.44 million, the first decline in five years, though still historically high. A more telling detail buried in the same report: organizations with widespread unauthorized AI use, so-called “shadow AI,” paid an extra $670,000 per breach on average. Among organizations that had suffered an AI-related security incident, 97% lacked proper access controls, and 63% had no AI governance policy at all.

The US number sharpens the point. American organizations paid an average of $10.22 million per breach in 2025, up 9% year over year and the highest of any country IBM tracks, driven in large part by steeper regulatory fines. When ownership and provenance aren’t clear, the risk doesn’t disappear. It compounds quietly, until it lands as a seven-figure number on someone’s desk.

Driver Three: Regulators Are Catching Up

If the first two forces come from the market and from risk, the third comes from institutions moving on their own. In the EU, the data economy was valued at nearly €325 billion in 2019, about 2.6% of GDP, and the European Commission’s own Data Market Monitoring Tool projects it will pass €550 billion by 2025, close to 4% of the bloc’s GDP.

The regulatory scaffolding is catching up to match. The EU Data Act became applicable in September 2025, giving users and qualifying third parties new rights to access the data generated by connected devices and cloud services, a direct legislative push toward treating data as something meant to be shared, priced, and moved, rather than something a single platform can quietly sit on.

Markets and institutions are, for once, pulling in the same direction.

From Concept to Real Market

Growth rates on their own can feel abstract. Actual transaction volume is more convincing.

Start with the model data monetization has run on for two decades, the data broker industry. It’s already worth an estimated $433.9 billion in 2025, projected to reach $616.5 billion by 2030, growing at roughly 7.3% a year. That’s a genuinely enormous market, and almost none of that value flows back to the people whose data is actually being bought and sold. It flows to the intermediaries sitting between the data and its source.

That gap, between the scale of the market and who actually gets paid, is exactly what data assetization is trying to close. The category isn’t a whitepaper thought experiment anymore. Hundreds of billions of dollars are already moving through data commerce every year. The open question isn’t whether data has value. It’s who captures it.

Two Paths Running in Parallel

Right now, the race is being run down two roads at once.

One is top-down and institutional: regulation like the EU Data Act, formal registries, and compliance-driven data-sharing frameworks that fold data assets into existing market oversight. Its strength is legitimacy and scale, built for enterprise-to-enterprise and industry-level data flow, and it plugs cleanly into systems regulators, auditors, and large companies already trust.

The other is bottom-up and technical: cryptography, on-chain identity, and programmable asset protocols that try to make ownership something established automatically the moment data is created, rather than something a central authority has to approve case by case. Smart contracts then track and route the resulting revenue back to the owner every time that data gets used. The imagination here runs toward two scenarios the institutional path struggles to reach, personal data assetization at the individual level, and the high-frequency, small-value settlement an AI-agent economy is going to need. A single post, a browsing history, an individual creator’s back catalog, none of it can realistically go through an enterprise-grade registration process, but all of it is exactly the kind of thing an AI agent might want to call, and pay for, thousands of times a day.

Traditional capital has started paying attention to this second path too. Earlier this year, venture firm a16z crypto published a report titled “AI Needs Crypto — Especially Now,” arguing that as AI’s ability to impersonate people improves, the right response is to pull identity verification out of centralized platforms entirely and build a verifiable, portable, on-chain identity layer instead, paired with blockchain’s ability to handle the high-frequency micropayments AI agents will need to transact with each other. Top-down policy design around data-market formalization, and bottom-up technical work on data-ownership protocols, are converging on the same destination from two very different starting points.

Still a Few Miles From Maturity

None of this means the category is finished. A few real gaps remain.

Standards haven’t converged. How data assets get registered, verified, and priced is still being worked out through parallel, competing approaches, both in policy circles and in code, nothing close to the kind of universal rulebook that exists for stocks or bonds.

Supply and demand don’t line up cleanly yet. Not all data is valuable by default. What AI actually pays for is structured, high-quality, verifiable data, and most of the raw data individuals and companies have sitting around isn’t there yet. That gap is exactly why labeling, cleaning, and usability verification have quietly become some of the fastest-growing segments in the whole pipeline.

And circulation has to find a balance between privacy and monetization. That tension is precisely why “usable but invisible” privacy technologies, zero-knowledge proofs, trusted execution environments, are moving out of research papers and into production faster than almost anyone expected. How well that transition goes may end up being the single biggest variable in whether data assetization ever reaches real scale.

Where This Leaves Us

Back to that opening number. The world is on track to generate something like 221 zettabytes of data in 2026 alone, up more than 20% from just a year earlier, a figure that’s already hard to hold in your head, and it keeps growing.

The data getting bigger isn’t in question. What’s still genuinely uncertain is how much of it ever clears the three hurdles, ownership, circulation, monetization, and turns from a silent byte into something with an owner, a price, and the ability to keep moving. Right now, most of the money still flows to the intermediaries standing between data and its creator, not to the creator.

The rules for this race are still being written, and the field keeps adding players, but the direction is already clear. In the next few years, the ability to turn data from something that merely exists into something that functions as an asset is going to become a real yardstick, for economies, for companies, and for every individual trying to hold their own in an AI-driven world.

Sources