This month I’m switching off Content Shield.

Not because it failed. Because we won.

The itch

The idea arrived before ChatGPT did. In 2022 I was booked to speak with various corporations about the AI text tools already on the market. The tools weren’t exactly great. But they set a predictable trend, plain to those of us carrying unhealthy levels of cynicism: machines that read the whole internet and write it back to you. Plagiarised, but laundered through synonyms and fresh sentence structure for the artifice of originality. No sources attached.

Destroying the financial incentive to publish, and demoralising the people doing the hard work at the very source: the collection of original insight.

I’d spent the best part of a decade in SEO, so I knew exactly how the reading happened: crawlers. And I could code, and I could market. It’s rare that a problem walks past wearing your exact measurements.

Then it got personal. People were plagiarising my journalistic side projects with automation, lifting articles the moment they were published, laundering them through AI, and republishing them with their ads running on them. My SEO cannibalised by my own words, at machine speed, stripping me of the bounty of my own work.

I called it what it was at the time: laundered plagiarism. Copy-and-paste with extra steps.

So Content Shield was a matter of principle above everything else, built in direct response to the people gutting my corner of the journalistic space. Somebody had to close the door. It might as well be the one with the crawler diagrams already on the whiteboard.

And underneath it all sat a sincere fear, on record at the time: that all the good stuff would vanish from the free web. That original work would retreat behind paywalls and log-ins, the open internet would clam up, and we’d be left with an echo chamber of machine summaries of machine summaries.

The first signs appeared almost immediately. Twitter killed its free API within weeks of ChatGPT arriving, and Reddit priced its own API into the stratosphere months later, taking a generation of third-party apps with it. Reddit’s CEO was refreshingly candid about why: they weren’t about to hand their “corpus of data” to the world’s largest AI companies for free. The clam-up had begun.

The product

So in December 2023 I built the door and closed it. Content Shield: a WordPress plugin that stopped AI crawlers reading your content, without touching your Google rankings. That last part was the hard part, and the part I’m still proudest of.

It worked. Ask ChatGPT to read a protected site and it returned an apology instead of your paragraphs.

The feature set grew with the fight:

  • Bot blocking. The core. AI crawlers simply couldn’t read the page.
  • Text locking. Ctrl-C is a threat vector too; this asked browsers to stop copy-to-clipboard.
  • Bot count. A running tally of how many scrapers we’d turned away. Weirdly satisfying reading.
  • A copyright notice in the site footer, telling any crawler that bothered to look that the content was not up for grabs.
  • Hands-on install support. Writers aren’t sysadmins, so I set it up for them personally.

My favourite came later: the Trap Street Mechanism. Cartographers used to draw a fake street into their maps; if a rival’s map showed the same street, they’d caught their copyist red-handed. We did the same with a tiny, harmless detail of misinformation planted in the copy, crossed with a bit of AltaVista-era black-hat: white text on a white background, invisible to the reader, dutifully hoovered up by the fledgling crawlers. If your “original” article contained our trap street, we had you.

An old cartographer’s trick and a nineties spam technique, repurposed as an evidence factory. I remain quite pleased with that one.

The beta went onto a healthy roster of third-party sites, mostly journalistic outfits and original-research publishers. Exactly the people with the most to lose. The interest was real, the product was live, and for a while I ran one of the very few tools in the world doing consent-based AI crawling.

The receipts

Two predictions from those launch weeks, both on the record.

First, from December 2023: “I would place a bet that within the next year, there will be an industry for Generative-AI-Crawler-Optimisation.” That industry now has a name (several, in fact: GEO, AEO, LLMO), conference tracks, and retainers.

Second, from the launch post and the royalties piece: that consent, not forgiveness, would have to become the default; and that the answer to “how would royalties even work?” was simple. They built a machine that writes like Shakespeare; they can figure out an invoice.

The market caught up

And then, right on cue, the corporates swooped in like the eagles at the end of The Lord of the Rings. Lovely of them to turn up, once the rest of us had walked the thing to Mordor on foot.

The Guardian, the BBC, and most major publishers now block AI crawlers at the door. Amazon polices bot traffic with the same logic. And in 2025 Cloudflare, which handles roughly a fifth of the web’s traffic, started blocking AI crawlers by default and attached a price list: pay per crawl, then pay per use. Permission over forgiveness, as infrastructure.

That was my product. As a default setting. On 20% of the internet.

I won’t claim a £10-a-month WordPress plugin moved Cloudflare’s roadmap. But grassroots pressure, and a working demonstration that consent-based crawling could be done without torching your rankings, is exactly how industries get nudged. We were early, we were loud, and we were right.

The door I built is now everyone’s front door. That’s not a loss. That’s the point.

The courts are catching up

The copyright argument I was making in 2023, that training on unlicensed work has to cost something, stopped being a blog post and started being a bill.

The New York Times case against OpenAI and Microsoft, the one I wrote about the week it landed, is still grinding through the system. But Anthropic has already paid: a $1.5 billion settlement, approved in July 2026, for training on pirated books. Roughly $3,000 per title. The largest copyright recovery in history.

And here’s the detail that tells you where we actually are: the same courts ruled that buying books, slicing off their bindings, scanning them, and binning the originals is transformative fair use. Anthropic destroyed millions of print books this way, rare and vintage editions among them, under the banner of training data. Digitise the library, skip the library fine, lose the library.

Piracy costs $1.5 billion. A guillotine and a scanner is fine. We are still very early in working out what consent means at machine scale.

I’ll confess, as I did back then: I’m no legal expert. But I hope the compromises the courts land on are healthy ones. Consent, credit, and a fair invoice. They built a machine that writes like Shakespeare; they can figure out the paperwork.

What I’d do differently

One thing. I bootstrapped the beta and stopped at the door marked “capital.”

The window between December 2023 and Cloudflare’s move was about eighteen months. With funding, a protocol instead of a plugin, and a team instead of a me, Content Shield could have been the standard rather than the proof of concept. I knew the wind; I under-bought sail.

Still. As an exercise in reading a market trend and shipping a real answer into it, while it still mattered, I’ll take it. UrbanFox taught me every side project deserves an ending written down. This is Content Shield’s.

The flag

Here’s the thing the whole saga proved, and it’s the thesis this entire site sits on.

Because all the while Content Shield was holding the door, the day job never stopped. I was working with clients whose content output was their livelihood, and others for whom content was the primary driver of traffic and the shop window for their expertise. Playing defence for them was never going to be enough; we needed an offence. After a lot of experimentation, we landed on new methodologies to publish, syndicate, and monetise their expertise: sense what the market is asking, draw the insight out of the experts who hold it, give it a human face, distribute it properly, and let the transactional traffic follow. That method is now the front page of this site.

Blocking the machines was only ever half the answer. The other half is the game now: being the source the machines, and the humans, have to cite.

Regurgitation is a done game. The winners will be the ones worth quoting.

Original qualitative and quantitative insight, with a human face on it. Research that doesn’t exist anywhere else. Expertise that’s visible enough to be named. If you’re not the citation, you’re the training data.

AI-laundered plagiarism is dead. Long live AI citation.