This week, it was announced that the New York Times is suing Microsoft, now owner of OpenAI (the company responsible for ChatGPT), to the tune of billions of dollars.

Although quite the Christmas stocking filler, this comes as not much of a shock. After all, it’s not like OpenAI trained ChatGPT on exclusively out-of-copyright Dickens novels and Creative Commons works.

I’ve personally felt the sting of AI plagiarism. It’s disheartening to see your hard graft copied, transformed, and paraded as new by AI tools and the people who publish the output as their own. I’ve made my feelings clear on this: it’s laundered plagiarism, a sophisticated way of copying and pasting that bypasses traditional checks and balances, not to mention it fundamentally disincentivises the open publication of unique written works.

Since the launch of Content Shield, I’ve been approached several times by writers who have observed their works plagiarised by ChatGPT and its users. “Let Disney’s lawyers deal with it” has seemingly become my catchphrase. I never thought I’d find myself advocating for Disney when it came to an issue of copyright; 2023 has truly been a wild ride.

The lawsuit has put a spotlight on the issue. The result of this case has the potential to be one of the most important legal precedents set in recent history, with untold and far-reaching consequences.

I am willing to be wrong and remain open to changing my mind at any time if presented with a compelling counter-argument. I am certainly no expert in law, but the following are my thoughts on the case, as understood on December 29th 2023.

I feel like I’ve heard this song before

I am not a lawyer, but I am a (former) musician. Not a good one, but that’s beside the point.

Just as ChatGPT has been ‘trained’ on copyrighted content without consent, so too has the music industry grappled with the complexities of using existing works to create something new. But there’s a critical difference: the music industry has established rules for sampling and interpolation, requiring licences and royalties to protect original creators. It’s high time we applied similar principles to text-based AI output.

While it’s tempting to treat generative AI utilising crawlers to scoop up swathes of the internet like piracy, it’s not quite that. And although the output functions as a replacement for the original work, it isn’t a direct copy.

Comparing AI to search engine indexing doesn’t quite fit either. There’s a fine line between helping navigate content and replacing it. (I have more to say about zero-click search results, but that’s for another time.)

The legal showdown may hinge on precedents from the music industry.

A whistle-stop tour of recorded music copyright

In the UK, recorded music copyright is a symphony of intricacies. When a track is recorded, there are two key rights: the rights to the music (composition) and the rights to the recording itself (performance).

If you cover a song, you’re reinterpreting the composition, needing a mechanical licence. You are responsible for the performance, so those royalties go to you; someone else owns the composition, so royalties go to them too. While you need a mechanical licence, you can record and publish covers of Taylor Swift songs to your heart’s content without asking Tay-Tay.

I have been paying Swift royalties since 2015 due to a published “We Are Never Getting Back Together” cover recording. I know she’d hate it.

Sampling, however, slices into both rights. It’s using a piece of the actual recording, requiring permission from the record label (performance rights) and the songwriter or publisher (composition rights).

Interpolating a song also treads on these dual paths, but in a nuanced manner, demanding a balance of respecting the original composition while introducing new performance elements. Each note played in this legal orchestra requires careful consideration to avoid discord.

Additionally, the public performance of a live or recorded song requires separate licensing by the venue or performer. And even more confusingly, there are different royalties for music and lyrics, but let’s not get bogged down in detail.

Of course, it goes without saying that piracy, creating unauthorised copies of a recorded work, is illegal. We all remember the greatest ad of a generation.

Are generative AIs sampling or performing?

If you’ve read this far, you can likely tell I believe generative AIs should pay royalties on the material from which they are ‘trained’ when used for commercial purposes. I am, however, conflicted about which type of royalties, who should pay (the users who republish its work or the generative AI company itself), and the mechanism by which they should be paid.

Until recently, I thought of generative AI as a cover version. However, two words caught my eye in the case: “verbatim excerpts.” If that holds, I see this as sampling unless correctly cited. Even in works published by a human being, there are limits to how much you can pull quotes from existing works under fair use.

My current thoughts

  • Generative AI should pay performance royalties, similar to mechanical licences for an unauthorised cover version. Think a fund, a bit like the PRS licence, that AI developers chip into, compensating the brains behind the original content.
  • Publishers who use AI should pay royalties similar to sampling, requiring consent from the original publisher. If you’re using AI to jazz up your content, you’re sampling in a way. And that means paying up and getting permission.
  • Original publishers should have the right to withdraw content from AI training sets. Sometimes they might not want their work to be part of this AI-driven future. And that’s okay. For this to work, AI developers need to be crystal clear about where they’re getting their data, and make it easy for creators to say “no thanks.”
  • Consent should be the golden rule in generative crawling: a permission-over-forgiveness model. Instead of AI developers hoovering up content willy-nilly, they should ask first. Think of it as knocking on the door before entering.

Where does this leave us?

As we stand at the crossroads of technological advancement and artistic integrity, it’s clear the path forward isn’t black and white. The lawsuit involving the New York Times, Microsoft, and OpenAI isn’t just a legal skirmish; it’s a clarion call for redefining the rules of the game in the age of AI.

Generative AI isn’t a pirate in the traditional sense. It’s more of a sampler, remixing and reimagining our collective digital tapestry. But, as with any art form that borrows from existing works, there’s a moral and legal obligation to acknowledge and compensate the original creators.

To me, the question isn’t whether AI should pay its dues. It’s how.

The outcome of the NYT lawsuit could be a watershed moment, setting a precedent that guides us through these uncharted waters. But regardless of the verdict, one thing’s for sure: the conversation around AI, copyright, and royalties isn’t just necessary; it’s overdue.

As we move forward, let’s keep sight of the core principle that has driven artistic creation for centuries: respect for the creator’s rights. It’s time to harmonise the tunes of technology and creativity, ensuring that as AI writes the next chapter of human innovation, it does so by playing a fair, respectful melody that honours those who laid the foundation for its symphony.

Oh, and if you have questions about the practicality of such a royalties mechanism: remember that they built a machine that can write like Shakespeare. They can figure it out.