AI literacy · Using AI well · Lesson 4 of 4
Checking outputs
Verify before you trust.
11 minute read
Language models produce fluent, confident text whether they are right or wrong. The tone never wavers. A correct date and an invented one arrive in the same calm, authoritative voice, which means the feeling of this sounds right is worth nothing as evidence. Checking outputs is not paranoia. It is just the workflow.
Why models get things wrong
To check outputs well it helps to understand why the errors happen, because they are not random glitches. A model is not looking facts up in a database. It is predicting, word by word, what text would plausibly come next, and it learned plausibility from patterns in its training data. Most of the time the most plausible continuation is also the true one, which is why models are right so often. But when the model does not have the fact, the machinery does not stop and say so. It keeps predicting, and what comes out is whatever a correct answer would most likely look like: a plausible date, a realistic name, a citation formatted exactly like a real one. The error is not a malfunction. It is the system working as designed, on a question it could not actually answer.
What to check first
The highest risk items are the specific, checkable details: names, numbers, dates, quotes and citations. Models are strong on the general shape of a topic and weakest on precise facts, because a detail that is slightly wrong still fits the pattern of plausible text. A quote can be almost right, attributed to the wrong person, or entirely invented, and it will read exactly like a real one.
Make it cite, then actually look
Ask the model to cite sources for its claims. Then do the part most people skip: open them. Check that the source exists, that it says what the model claims, and that it is the kind of source you would accept from a person. Models can fabricate very convincing citations, complete with real sounding authors and journals. A citation you have not opened is decoration, not evidence.
This is not a hypothetical risk. In 2023 a New York lawyer filed a court document that cited six earlier cases supporting his argument, all supplied by a chatbot, complete with names, dates and quotable passages. None of the six cases existed. The court checked, the fabrication came out, and the lawyer was fined and made international news. He had even asked the chatbot whether the cases were real, and it said yes. That last detail is the real lesson here: asking the model to verify itself is not verification, because the confirmation is generated by the same machinery that generated the mistake.
A worked check, start to finish
Suppose a model tells you that a particular Australian law changed in a particular year, and you want the claim for an assignment. The check runs in three steps and takes under five minutes. First, existence: search for the source the model named and confirm it is real, on a site you recognise, such as a government site ending in .gov.au. Second, content: open it and find the actual claim, because a real source that says something slightly different is the most common failure, and slightly different can flip your argument. Third, currency: check the date on the page, because a law described accurately as it stood five years ago may have changed since. If the claim survives all three steps, cite the source you opened, not the model. If it survives none of them, you have just saved yourself from building on sand.
Cross check and mind the cutoff
For anything that matters, run a quick search and see whether independent sources agree. One reliable source beats a page of unverified model output. And remember that every model has a knowledge cutoff: a date after which it knows nothing. Ask about recent events, prices, laws or versions of anything and it may answer from a world that no longer exists, without flagging it. If your question depends on now, check now somewhere that actually knows.
Newer does not mean checked
A misconception worth clearing up: each new model release is better, so surely by now the problem is fixed. Newer models do make fewer of these errors, and the improvement is real. But fewer is not none, and the errors that remain are harder to spot, because everything around them has become more polished and more plausible. A rare error inside highly convincing text is more dangerous than a common error inside clumsy text, not less, because your guard is down. The workflow does not change with the version number. Verify the checkable details, whatever the model.
Where this bites in a normal week
Most of the checking you will actually do is small. Before a statistic goes into an assignment, open a source. Before you repeat a surprising claim in the group chat, spend thirty seconds searching, because being the person who spread the false version costs more than the pause does. Before you act on anything about medication, tax, your pay rate at work or the law, find the real source, which for pay in Australia means an official body like the Fair Work Ombudsman rather than a chat window. The pattern is the same each time: the model gets you to the neighbourhood fast, and the last hundred metres, the checking, is yours.
Calibrate, do not distrust everything
None of this means model outputs are useless. Most of what they say about well documented topics is broadly right. The skill is calibration: match the checking to the stakes. Brainstorming needs almost none. A fact going into an assignment needs its details verified. Anything touching health, money or law needs a real source before you act. Confident tone is decoration. Treat it that way and the tool becomes safe to lean on.
This lesson closes the loop on the whole topic. Prompting gets you a useful answer, the sparring habit keeps the thinking yours, the study line keeps the learning real, and verification is the step that makes all of it safe to rely on. Skip it and everything upstream is resting on trust in a machine that cannot tell you when it is guessing.
Check your understanding
8 questions. Pick an answer for each, then check.
1. Why is a confident tone in a model's answer not evidence it is correct?
2. Which details in an output carry the highest risk of being wrong?
3. After asking a model to cite sources, the essential next step is to
4. A knowledge cutoff means a model
5. The lesson's advice on how much to verify is to
6. In the 2023 New York lawyer case, why did asking the chatbot whether its cited cases were real not help?
7. In the worked check, why is reading the actual claim a separate step from confirming the source exists?
8. The lesson says errors from newer, more capable models are