Curiosity

AI literacy · Ethics · Case study

The screening tool

A fictional company, a real pattern. Work through what the tool did to Hannah's application, and what everyone should have done.

Every November, a large Australian retail chain hires thousands of casuals for Christmas. This year over 40,000 applications arrive for around 6,000 positions, and for the first time the company uses an AI screening tool to rank them. Recruiters will only read the top slice. The tool was trained on ten years of the chain's own hiring records: who was hired, and who among them stayed and was rated well by store managers.

Hannah is 17 and in year 11. Her application is genuinely strong: eighteen months at a busy local bakery, a coach's reference from netball, and availability across the whole Christmas period, which stores desperately want. The tool ranks her in the bottom third. No human ever reads her application. She receives a polite automated rejection that mentions a high volume of strong candidates, and she concludes she must have done something wrong.

Priya, a data analyst on the company's people team, is asked to prepare a routine report on the new tool. Checking the rankings, she notices something odd: applicants from a cluster of suburbs and schools score consistently high, while strong applications from elsewhere sink. Digging in, she finds why. For ten years, the chain's stores mostly hired from the suburbs near their biggest locations, often through word of mouth from existing staff. The model learned that pattern faithfully. It is not weighing bakery experience much at all. It has learned what past hires looked like, and past hires looked like the suburbs the company already knew.

There is a second problem hiding under the first, and it is worth pausing on because Hannah never sees it. If the tool runs unchanged for a few seasons, its bias does not stay level. It compounds. The applicants it ranks highly get hired, work their seasons and become the next round of training data, proving the tool right, while the strong applicants it buried never get the chance to generate any record at all. The company will never learn what Hannah would have been like as an employee, so the data can never correct itself. Within a few years the tool would not just be reflecting the old hiring pattern. It would be enforcing it, with each season's results laundered into fresh evidence.

Priya raises it. The debate inside the company is genuinely hard, and both sides have a point. The recruitment lead is blunt: the tool saved roughly 300 hours of manual screening this season, the team cannot read 40,000 applications by hand, and the tool is only reflecting our own hiring history, which nobody complained about at the time. Priya's side is equally blunt: that history is exactly the problem, the tool is now applying it to 40,000 people at once with a precision no biased manager ever managed, and a 17 year old with a strong application never got a human glance because of her postcode.

The company lands on a middle path. It keeps the tool, but changes its job: instead of producing one ranked list, the tool now sorts applications by availability and basic requirements only, screening out only those missing hard essentials like working rights. Recruiters read a random sample from every region, and the company commissions an external audit of the tool's scores against outcomes for different groups before next season. It also rewrites the rejection email to say that automated screening was used and to name a contact for applicants who believe their application was not fairly assessed.

The external audit is not a vague health check. It has specific questions to answer. Do applicants from different regions and schools with similar experience receive similar scores? Do the tool's decisions actually predict anything about later job performance, or was it only ever predicting resemblance to past hires? And does the random human sample keep surfacing strong candidates the automated sort would have buried, which is both a check on the machine and a measure of what the old process was silently costing? If that audit had existed before launch, Priya's discovery would have been a pre launch finding instead of a mid season scramble, and before deployment is always the cheapest place to catch this class of problem.

And Hannah? Under Australian law, discrimination in hiring is unlawful whether a human or a machine does it, but in practice her options were thin. She never knew a tool was involved, so she had nothing to question. The redesigned process at least tells applicants that automation was used and gives them a door to knock on. This season, a friend who works at the chain encourages her to reapply. Her application, read this time by a human recruiter in the regional sample, gets her an interview inside a week.

Your tasks

Work through these in order, on paper or in a doc. They are the point of the story.

  1. 1List everything that went wrong before Priya's report, and at what stage each problem entered: the training data, the model's design, the deployment or the rejection email.
  2. 2Argue the recruitment lead's side as strongly as you can in a short paragraph: 40,000 applications, 300 saved hours, and a tool that only mirrors the company's own history. Then argue Priya's side just as strongly. Do not strawman either one.
  3. 3Explain why the tool ranked Hannah low even though it was never given her address as a scoring rule. What did it actually learn from ten years of hiring records?
  4. 4The bias lesson described a feedback loop. Explain in a few sentences how this tool would have created one if it had run unchanged for five seasons, and why the missing data about rejected applicants means the loop could never fix itself.
  5. 5The company's fix keeps the time savings but changes the tool's job. Design your own alternative process that screens 40,000 applications fairly without giving up the 300 hours. Specify exactly what the machine does and what humans do.
  6. 6Hannah had almost no practical recourse at first. Write the three things you think any applicant judged by a machine should be entitled to, and note who would have to pay for each one.
  7. 7The original rejection email said nothing about automated screening. Rewrite it honestly in four sentences or fewer, including what applicants can do if they believe the assessment was unfair.