AI literacy · The key terms · Case study
The headline, decoded
A news headline says AI beats doctors. Maya uses four lessons of vocabulary to find out what it really says.
Maya is 17 and wants to study medicine. So when a headline crosses her feed, New AI outperforms doctors at diagnosis, she stops scrolling. Her first instinct is excitement. Her second, fresh from this topic, is a list of questions. She opens the article and starts pulling the claim apart, word by word.
What model? The article names it: SkinCheck Pro, built by a startup called Dermalytic, a model trained to classify photos of skin conditions into 14 categories. So not doctors in general, and not diagnosis in general. One model, one narrow task: matching photos to category labels. A doctor's diagnosis involves touching the skin, asking about history, and noticing the thing the patient did not mention. The model does none of that. The headline said diagnosis. The study measured photo classification.
Trained on what data? Maya finds the detail in paragraph nine: 130,000 photos collected between 2016 and 2024 from two teaching hospitals in one European country. She thinks about missing data. Two hospitals, one country, means one mix of patients, and she knows from the bias lesson that skin conditions can look very different on darker skin. If darker skin is rare in those 130,000 photos, the model has barely seen the thing it would be asked to judge in a country like Australia.
Tested how, against whom? The test set was 800 photos held back from the same two hospitals. The model got 91% right. The comparison group was 9 doctors, a mix of dermatologists and general practitioners, who scored 84% on average. But the doctors were shown only the photos: no patient history, no questions, no examination. Maya writes in her notes: they did not test doctors doing their job, they tested doctors playing the model's game, on the model's home ground. The dermatologists alone scored 89%, a gap of two photos in a hundred, which the headline rounded up to outperforms doctors.
Who ran the test? Near the bottom: the study was funded by Dermalytic, and four of its six authors work there. That does not make the numbers fake, Maya reminds herself. It does mean the people measuring had every reason to measure generously, and that she should want to see someone independent repeat the result before trusting it.
Then she finds the paragraph the headline ignored. As a follow up, the researchers ran the model on 500 photos from a clinic in a different country, with different cameras, different lighting and a different mix of patients and skin tones. Accuracy fell from 91% to 76%, and it fell furthest on darker skin. On data unlike its training data, the model lost most of its advantage. That is the sentence Maya underlines twice, because it is the one that describes the real world, where patients do not arrive matching the training set.
She also writes down what would change her mind, because that is the fair test of whether she is being sceptical or just cynical. An independent team repeating the study on new patients would help. So would a trial on photos from Australian clinics, with results reported separately by skin tone, so the follow up problem gets measured rather than hidden inside an average. And if the tool were ever pointed at real patients here, it would need to satisfy the Therapeutic Goods Administration, the Australian regulator that treats diagnostic software as a medical device. A headline cannot grant that approval. Evidence can.
The last thing Maya notices is how the framing shaped her feelings. Outperforms doctors made the model sound like a rival, and for a moment had her doubting the career she wants. Helps doctors look twice describes the same technology as a colleague, a second set of eyes that never gets tired at the end of a clinic day but also never asks the question that cracks a case open. The numbers support the second framing far better than the first, since the model barely edged out the specialists on its home ground and fell well behind everywhere else. Which framing sells more clicks is a different question, and it explains a lot of AI coverage.
Maya closes the article with a very different summary from the headline. One model, trained on 130,000 photos from two hospitals, beat 9 doctors at photo classification on photos from those same hospitals, in a test its maker funded, and slipped badly on photos from anywhere else. Still genuinely interesting. Possibly even useful one day, as a tool that helps doctors look twice. But it is a smaller, more specific and more fragile fact than outperforms doctors at diagnosis. The vocabulary did the work: model, training data, test set, benchmark, missing data, bias. Six terms turned a headline into a set of answerable questions.
Your tasks
Work through these in order, on paper or in a doc. They are the point of the story.
- 1Find a real news headline from the past year claiming an AI beats humans at something. Save the link, and note where you found it.
- 2Identify the model behind the headline: who built it, what specific task it performs, and how that task differs from the human job named in the headline.
- 3Find what the model was trained on. If the article does not say, note that silence as your first finding, and write down who is most likely to be missing from that data.
- 4Describe the test: what benchmark or test set was used, who the human comparison group was, and whether the humans were doing their real job or a narrowed version of it.
- 5Follow the money: who ran and funded the evaluation, and has anyone independent repeated it?
- 6Rewrite the headline so it says only what the evidence supports, the way Maya did. Compare your version with the original and list what the original left out.
- 7Write down the one follow up test that would most change your mind about the claim, the way the 500 photo follow up changed Maya's reading, and note whether the article says anyone has run it.