Cloudflare Locks Out AI Crawlers — MEMO Lets Agents Find Their Own Data

Cloudflare Locks Out AI Crawlers — MEMO Lets Agents Find Their Own Data

On September 15, Cloudflare is rolling out a change: for every newly onboarded website, if the page carries advertising, Training and Agent crawlers will be blocked by default, while Search crawlers continue to be allowed through.

At first glance, this looks like a routine product policy tweak. Look closer, and it draws a line: the era of freely, unauthorized, mass-scraping the web to feed AI is being structurally shut down.

Why is the Search crawler exempt? Because it serves humans — it’s a traffic gateway that brings a site more visitors, and site owners are happy to leave that door open. Training and Agent crawlers serve machines — they lift page content wholesale, without sending traffic back or paying anything in return, so of course site owners won’t leave that door open for them. What Cloudflare is really doing here is separating “what humans want to view” from “what machines want to take,” pricing and authorizing each one separately.

Once this door becomes the default setting, the AI industry runs into a real problem: where does the data come from now?

The Old Playbook Is Breaking Down

For the past decade, the default way AI got its data was scraping. Whoever’s crawler ran fastest and widest ended up with the biggest dataset. That logic only worked because website owners never had the time or the tools to distinguish a human visitor from an AI crawler.

Cloudflare covers more than a fifth of global web traffic, and the moment it flips that default, scraping itself doesn’t disappear — but its default legitimacy does. Every newly onboarded site starts closed from day one. AI companies now have three options: negotiate a license, pay for access, or get nothing at all.

It’s easy to predict that other content owners beyond Cloudflare will come to the same realization: data has an owner, and it can’t keep being taken for free.

What MEMO Has Been Building All Along Is the Answer to This Problem

One thing worth being clear about upfront: MEMO isn’t here to help anyone sneak around Cloudflare to keep scraping. The real opportunity isn’t in “getting around the lockout” — it’s that once the free-riding path closes, the industry needs a path that was already rights-confirmed, already compliant, and already built around paying for data. MEMO has been laying that path for years.

Data flowing through MEMO’s storage system falls into three categories from the source, and each one comes with authorization built in.

Data users publicly share. Once a user uploads data into the MEMO storage system, whether to make it public, and to whom, remains entirely the user’s own decision — they can authorize access to a specific party, or choose to make it fully public. From the very first step of upload, ownership is unambiguous, and whether the data goes public is a decision the user makes themselves — it can’t be scraped away quietly by someone bypassing consent.

Data tokenized through ERC-7829. ERC-7829 is MEMO’s proposed data asset NFT standard, packaging heterogeneous content — documents, datasets, AI interaction logs — into on-chain assets that can be owned, priced, and traded. Every time the data is used, revenue automatically flows back to its owner. This is the opposite of free-riding: use it, and you pay; pay, and the split happens automatically.

Open data in the data marketplace. The data asset platform connects uploading, minting, management, and trading into one complete loop. Buyers and sellers settle through smart contracts — whoever wants to use the data pays for it, prices are discovered by the market, and no middleman is needed.

All three sources point to the same thing: data isn’t taken — it’s traded. What Cloudflare is tightening is unauthorized scraping. What MEMO has always offered is authorized circulation. That’s not a coincidence — it’s the same industry problem being approached from both ends, meeting in the middle.

Getting the Data Is Just the Entry Point — The Real Question Is How Machines Run This Whole Process on Their Own

But if the story stops at “MEMO has a clean source of data,” it’s only half told. Where the data comes from is the first-layer question. The second layer is harder: when an Agent goes out to find data on its own, negotiate a price on its own, and complete the transaction on its own, what does it use to prove who it is, what does it use to pay, and who owns the new data it generates along the way?

These three questions aren’t three isolated needs — they’re a causal chain. Miss one link, and nothing before it can run.

Step one, the identity layer — an Agent first has to be able to prove who it is. DID assigns a unique decentralized identity marker to every user and every piece of data. ERC-8004, which MEMO has integrated, is an identity and reputation standard purpose-built for autonomously operating AI Agents. Before an Agent can access rights-confirmed data, the first thing it has to do is present a queryable, verifiable on-chain record — what it’s done, whether it’s defaulted on anything, what its reputation score is. Without this step, a data owner has no reason to grant access to an anonymous black-box program.

Step two, the payment layer — once identity is verified, payment has to settle on the spot. The x402 protocol, integrated by MEMO, makes payment as simple as a single API call, with granularity fine enough for a single request or a single chunk of data. The Agent economy is naturally high-frequency and low-value per transaction — a single call might be worth only a few cents, but the volume of calls is enormous, and traditional payment methods’ fees and settlement cycles simply can’t keep up. x402 solves exactly this bottleneck: cash and goods change hands atomically and simultaneously, with no credit risk involved.

Step three, data asset formation — the new data an Agent produces along the way can’t just be used once and thrown away either. As an Agent calls on data, executes tasks, and produces new interaction records and accumulated knowledge, that output can likewise be packaged into an asset through ERC-7829 and stored in MEFS, ready for the next call or the next Agent to use. This step turns the entire chain from a one-way extraction into a regenerating loop — data gets used and, at the same time, generates new data that can itself be rights-confirmed.

Stack these three layers together, and you get one complete pathway: identity lets an Agent in the door, payment lets a transaction settle on the spot, and asset formation lets an Agent’s own output re-enter the next round of circulation. This isn’t three features bolted together — it’s a single chain held up by one integrated system, and getting the data is just the first link of that chain breaking the surface.

One Step Further

If this chain can scale, a bigger narrative naturally grows out of it: AI stops depending on one-time scraped static datasets, and instead continuously and autonomously acquires data through identity verification and real-time payment, while feeding its own output back into the same system — which already looks a lot like the early shape of “AI training and evolving on its own.” That’s a much bigger topic, and one worth its own separate piece down the road.

Back to the present: after September 15, the default permission for free-riding scraping is being shut off, one site at a time. MEMO has no intention of fighting this trend, because the direction has always been the same one MEMO has been arguing for years — data sovereignty belongs back with the people who create the data. Cloudflare has just proven it for us again: this direction is the right one.