Curiosity

AI literacy · Responsible use, and what comes next · Lesson 3 of 4

Reading AI claims critically

Hype, benchmarks and headlines.

10 minute read

Every week brings a new AI headline. A model has passed some famous exam, or will replace some profession, or threatens civilisation itself. Most of these claims are neither lies nor truths. They are marketing, compressed until the caveats fall off. Reading them well is a skill, and it is the same skill you would use on any other claim someone profits from your believing.

Marketing versus evidence

A company announcing its own model's results is marketing, even when every number in it is accurate. Evidence looks different: independent researchers testing the claim, results that hold up when others repeat them, and details about what was measured and what went wrong. The gap between a launch demo and independent testing is where most AI disappointment lives. Extraordinary claims need extraordinary evidence, and a company's own highlight reel is ordinary evidence at best.

Why does independence matter so much? Because of what never gets shown. A launch demo is chosen from many runs, and you see the best one, never the seventeen attempts that failed before it. That is not necessarily lying. It is selection, and selection quietly converts an unreliable system into an impressive one. Independent testers have the opposite incentive, because they earn attention by finding the cracks. So a claim that survives people actively trying to break it has passed a real test, and a claim that has only survived its own launch video has passed a rehearsal.

What benchmarks measure, and what they miss

A benchmark is a fixed set of test questions, and a benchmark score tells you exactly one thing: how the model did on those questions. It does not tell you the model can do the job the exam was designed to filter humans for. A model that scores well on a medical exam has matched patterns in medical exam questions. That is genuinely impressive, and it is not the same as safely treating a patient. Benchmarks can also leak: if similar questions appeared in the training data, a high score partly measures memory, not ability. And a polished demo is a benchmark of one, showing the best run and none of the failures behind it.

How a careful claim becomes a wild one

Follow one invented claim down the pipeline and watch the caveats fall off. A research team reports that a model answered 82% of past law exam questions correctly, multiple choice sections only, using questions from years that may overlap its training data. The company announcement compresses that to: our model passes law exams. A tech site compresses again: AI now outperforms law students. By the time it reaches your feed it is a fifteen second video declaring that lawyers are finished. Nobody in that chain told an outright lie. Each step simply dropped a caveat, and each step had a reason to, because the company wants investment, the site wants clicks and the account wants followers. The original finding, a genuinely interesting one, is four compressions behind you by the time you meet it.

This is also why seeing a claim everywhere is weaker evidence than it feels. Most outlets do not retest anything. They rewrite the same announcement, so twenty articles can be one press release wearing twenty mastheads. Australian feeds are especially prone to this, because so much of our technology coverage is republished from overseas outlets with another caveat or two trimmed for length along the way. Counting sources only works when the sources are actually independent, and the checklist below asks about independence for exactly this reason.

Hype and doom are both for sale

AI will fix everything and AI will destroy everything are both excellent business models. Hype raises money and sells products. Doom sells clicks, books and attention. Both get amplified because extreme claims travel further than careful ones. The boring truth usually sits in the middle: a genuinely powerful, genuinely flawed technology, changing work and study in real but uneven ways. Boring rarely trends. It is usually what happened.

Where you will actually meet these claims

You will rarely meet these claims in a research paper. You will meet them as a thumbnail while you are half watching something else, as a post shared into the family group chat by a worried relative, or as a confident line in a classmate's presentation that was found in thirty seconds of searching. The checklist works in all three places, however the family group chat deserves a special mention, because you are now often the most AI literate person in it. The useful move there is not mocking the shared article. It is finding the boring version and offering it: here is what was probably demonstrated, and here is what got added on the way to the headline. That lands better than an eye roll, and every time you do it the skill gets faster.

The checklist

  • Who benefits if I believe this claim?
  • Is this independent evidence, or the claimant's own demo and numbers?
  • What exactly was measured, and is that the same as the ability being claimed?
  • Are the failures shown, or only the best runs?
  • What is the boring version of this claim, and would it still be a story?

Notice that this is the disclosure skill from the last lesson pointed outward. Disclosure asked you to describe your own AI use precisely, naming what the tool did and where the boundary of your work sat. Reading claims critically asks whether a company has done the same, and the ones that deserve more trust are the ones that publish their failures, name their test conditions and state the boundary of what was actually shown. The final lesson takes the last step: given a technology this hyped and this real at the same time, what is actually worth building in yourself?

Check your understanding

8 questions. Pick an answer for each, then check.

  1. 1. A company publishing impressive results for its own new model is best treated as

  2. 2. A model scores highly on a medical exam benchmark. This tells you the model

  3. 3. Benchmark leakage is a problem because

  4. 4. Why do both hype and doom dominate AI headlines?

  5. 5. The point of writing the boring version of an AI claim is to

  6. 6. Why can a launch demo mislead even when nothing in it is false?

  7. 7. In the invented law exam example, how did a careful finding become a wild headline?

  8. 8. Twenty articles all reporting the same AI claim count as strong evidence only if