Curiosity

AI literacy · Ethics · Lesson 2 of 4

Whose work trained it?

Copyright, creators and consent.

10 minute read

Every impressive AI model learned from an enormous pile of human work: books, news articles, artworks, songs, photographs and code, mostly scraped from the internet. The overwhelming majority of the people who made that work were never asked, never told and never paid. Whether that was fair, and whether it was legal, are two of the most fiercely contested questions in AI right now.

What training actually does with a work

Before weighing the arguments, it helps to be precise about what happens during training, because both sides lean on it. The model reads a work and adjusts millions of internal numbers so that it becomes slightly better at predicting what comes next in text like it. It does not file the book away in a library it can open later. What survives is a statistical residue: the rhythms, structures and associations of everything it read, blended together. That is why developers reach for the reading analogy. However the analogy is not perfect, because researchers have shown that models can sometimes reproduce passages very close to specific training texts, especially famous ones that appeared many times in the data, and an artist's distinctive style can be summoned by name. Something of the original clearly persists. How much persists, and whether that counts as copying, is precisely what the courts are being asked to decide.

The creators' case

Artists, authors, musicians and journalists argue something simple: you built a commercial product out of our work, without permission, and that product now competes with us. An illustrator's decades of style can be imitated in seconds by a model trained on her portfolio. A news organisation's reporting, which cost real money and sometimes real risk to produce, can be summarised by a chatbot that sends no readers and no revenue back. If anyone else copied work at this scale for profit, they say, we would call it theft. Calling it training does not change what happened.

The developers' case

AI companies argue that training is more like reading than copying. A model does not store the books it trained on; it learns statistical patterns from them, the way a human writer learns from everything they have ever read, and copyright has never covered learning from a work, only reproducing it. They also argue the practical point: no useful model could exist if every one of billions of documents required an individual licence, so demanding that means demanding these tools never exist, along with everything good they can do. This is a genuinely strong argument, which is why courts are taking it seriously.

Where the fight stands

Authors, artists, music companies and news organisations have filed major lawsuits against AI developers in several countries, and governments are reviewing their copyright laws. These cases are working their way through courts, and different courts may reach different conclusions, so be suspicious of anyone who tells you the law is already settled. In the meantime, some AI companies have started signing licensing deals with publishers, which quietly concedes that the work has value worth paying for.

Online does not mean free

One idea muddies this whole debate and is worth clearing away: the belief that anything posted on the internet is public domain. It is not. Copyright applies automatically the moment a work is created, in Australia and in most of the world, with no registration and no copyright symbol required. The drawing you post, the story you upload and the song you share are all copyrighted works, and posting them publicly gives people permission to look, not permission to reuse. So the question in these lawsuits was never whether the works were protected. They were. The question is whether training a model on a protected work is one of the uses copyright forbids, and that is a genuinely open question, which is different from an already answered one.

How Australia is different

The headline lawsuits are mostly American, and the American ones lean on a defence called fair use, a flexible test that lets courts allow unlicensed uses they judge to be fair. Australia does not have that defence. Australian copyright law allows only fair dealing, a narrower set of specific permitted purposes such as research, criticism, review, parody and news reporting, and training a commercial AI model does not fit neatly into any of them. So the same training run could be judged differently in different countries, which is one reason the global picture will stay messy for years. It also means Australian creators, from novelists to news outlets to the artist selling prints at a local market, are watching decisions in foreign courts that will shape what happens to their work here.

What a fairer deal could look like

Between free for all and total shutdown there is a wide middle: opt out registries that models must respect, opt in licensing with payment, collective schemes like the ones that pay musicians when radio plays their songs, and clear labels on AI outputs that imitate a living artist's style. If you make things, and most students do, this is not an abstract debate. It is about whether the next generation of creative work is something people can still get paid for.

The collective scheme idea deserves a closer look, because it already works elsewhere. When a cafe plays music or a radio station broadcasts a song in Australia, nobody phones each songwriter for permission. The venue buys a blanket licence from a collecting society, which pools the fees and distributes them to the artists whose work was used. The system is imperfect and artists argue about the split, but it solved the same structural problem AI training has: a use involving far too many works for individual deals. Whether something similar can be built for training data, who would run it, and how you would even measure whose work contributed what to a model's output are open questions. But the shape of a workable answer has existed for decades.

Your own work is in this story

And this reaches your own work sooner than you might think. If you post art, writing, music or videos anywhere public, they are exactly the kind of material scraping tools collect, and several major platforms have updated their terms of service to permit user content to be used for AI training, sometimes with an opt out buried in the settings and sometimes without one. It is worth actually reading what you agreed to on the platforms you use. It also cuts the other way: when you generate an image in the style of a working artist, you are standing on one side of this debate whether you meant to or not. Neither use makes you a villain, but both make you a participant, and a participant should know the argument.

This lesson and the next are really the same question from two angles: what happens when a machine is built from human work and then competes with the humans. Here the work is creative and the competition is direct. In the jobs lesson the work is everyday labour and the competition is quieter, but the pattern to watch is identical: who bears the cost, and who captures the gain.

Check your understanding

8 questions. Pick an answer for each, then check.

  1. 1. Most of the creative work used to train large AI models was

  2. 2. The strongest version of the developers' argument is that training is

  3. 3. The creators' strongest complaint is that the models

  4. 4. What complicates the developers' claim that a model is a reader, not a copier?

  5. 5. How does Australian copyright law differ from the American law the big lawsuits rely on?

  6. 6. Which statement about work posted publicly on the internet is true?

  7. 7. What is the current legal status of the major training data lawsuits?

  8. 8. What do licensing deals between AI companies and publishers quietly concede?