Plainstart Back to the kit

Article

How to Tell Whether Your AI Screening Tool Is Filtering People Out Unfairly

AI screening tools save time. They also carry a specific risk that most governance conversations miss: they can filter people out in patterns that have nothing to do with whether those people can do the job. This article is about how to spot that, how to question a vendor about it, and how to keep records that let you explain any decision later.

The data-leak risk gets most of the attention. This one deserves more.


Why Screening Bias Is a Different Kind of Problem

Most AI governance discussion for recruitment agencies focuses on data: where candidate data is stored, who can access it, how long it is kept, and what the privacy rules in your jurisdiction say about that. Those are real questions. They have mostly practical answers.

Bias in screening works differently. It does not announce itself. There is no breach notification, no audit log showing something went wrong. The tool just keeps running, and a pattern builds up quietly in your pipeline. Certain kinds of candidates stop making it through. You do not notice because the shortlists look reasonable and the clients are not complaining.

The exposure here is not a data problem. It is a judgment problem. The tool is making or shaping judgments about people, and those judgments may be systematically skewed in ways that have no connection to job performance.

Where does the skew come from? Usually from the data the tool was trained on. If a model learned from historical hiring decisions, it learned from human decisions, and human decisions carry patterns. If your agency, or the vendor's broader training set, historically placed more candidates of a certain background into certain roles, the model may have absorbed that as a signal of suitability. It is not doing anything strange. It is doing exactly what it was trained to do. The problem is what it was trained to do.

This is why the standard response of "our tool is just matching skills to requirements" is not sufficient. Skills-matching sounds neutral. In practice, how skills are described, weighted, and interpreted in natural-language processing is not neutral. The framing of a CV, the school someone attended, gaps in employment history, even the way a candidate phrases a sentence can all feed into a score in ways that correlate with characteristics unrelated to the job.


Questions to Put to a Vendor

Before you rely on any AI screening or ranking tool, ask these questions directly. Write them in an email so you have the answers on record.

What data was the model trained on, and was it audited for bias before deployment?

A credible vendor can describe the training data in general terms and confirm that bias testing was part of the build process. A non-answer sounds like: "Our model uses advanced machine learning to match the best candidates." That tells you nothing about the training data. Press again.

Can you show me outcome data broken down by demographic group?

You want to see whether the tool's accept or progress rates differ across gender, age band, or any other characteristic the vendor can report on. Some vendors will say they do not collect this because they do not ask candidates for demographic data. That is a reasonable constraint. Ask what proxy testing they have done instead, because proxy testing is how bias is found when direct data is not available.

What happens when the model encounters a CV that looks unusual?

Unusual means career changers, candidates with gaps, candidates from non-traditional educational backgrounds, candidates whose CVs are written in a second language. Ask for specific examples. A vendor who has thought about this can give you specific answers. A vendor who has not will say something general about the model being "flexible" or "context-aware."

Is there any way for a human to see why a candidate was ranked where they were?

This is an explainability question. If the answer is no, or if the explanation is a score with no reasoning behind it, that is a problem. You need to be able to explain a screening decision to a candidate or a client if asked. A black-box score does not let you do that.

What is your process when a client or candidate raises a concern about a decision?

A vendor with a real answer has a process. A vendor without one will improvise an answer that sounds like a process. Ask for it in writing.


How to Test the Tool Yourself

You do not need a data scientist. You need a spreadsheet, some historic data, and a few hours.

Pull the last 200 to 300 applications your agency processed for a role type you fill regularly. You need the original CVs or application data, the score or ranking the tool gave, and the outcome: whether the candidate was shortlisted, interviewed, or placed.

Now look for patterns in who did not make it through.

Sort by score and look at the bottom third. Read twenty of those CVs. Are there characteristics that appear repeatedly? Career gaps. Certain kinds of institution. Certain ways of describing experience. Non-standard career paths. You are not running a statistical test. You are reading for patterns that a reasonable person would notice.

Then look at the top third and read twenty of those too. What do they have in common that the bottom third does not? If the answer is mostly about how the CV is written rather than what the person has actually done, that is worth examining.

A Worked Example

A mid-sized recruitment agency specialising in office support roles decided to check their screening tool after one of their consultants noticed that placements had become less varied over eighteen months.

They pulled 240 applications for an administrator role type, all processed through the tool over the previous year. They sorted by tool score and read the bottom fifty CVs.

They found a pattern. Candidates who had taken a career break of more than eight months, for any reason, were scoring in the bottom quartile at a rate roughly double what you would expect if the tool were indifferent to gaps. The gap itself was not a stated criterion. It was not on the job description. But the tool had apparently learned something from the training data that made a gap a negative signal.

They then checked what happened to those candidates manually. Several had strong skills matches to the roles they had applied for. A few had been placed by the agency in earlier years, before the tool was introduced.

The agency went back to the vendor with this finding. The vendor confirmed that employment gaps did influence the model's scoring. They offered a configuration adjustment that reduced the weight of continuity signals.

The agency made three changes. They applied the configuration adjustment. They added a manual review step for any candidate scoring in the bottom third who had a gap of six months or more. And they started keeping a log of cases where a consultant overrode the tool score, with a brief note on why.

None of this required a data scientist. It required someone willing to read CVs and notice what they were seeing.


What to Record

Record enough to explain the decision later. For every candidate who is not progressed, you should be able to say what criteria were applied, what the tool returned, and whether a human reviewed that output.

If a consultant overrides the tool in either direction, that should be noted. If a candidate was progressed despite a low score, note why. If a candidate was not progressed despite a high score, note why. These notes do not need to be long. A sentence is enough.

The point is that "the tool said so" is not an explanation. A candidate, a client, or someone else may ask about a decision. You need to be able to answer.


Why the Human Review Has to Be Real

Many agencies add a human review step to satisfy a policy requirement. The consultant looks at the shortlist, approves it, and moves on. That is a rubber stamp. It is not a review.

A real human review means the reviewer has read the CVs, not just the scores. It means the reviewer is in a position to question the ranking and override it. It means the reviewer sometimes does override it.

If no one has ever overridden the tool, one of two things is true. Either the tool is perfect, which is not possible. Or the review is not real.

Build your process so that overrides are possible, recorded, and occasionally expected. A tool that is never questioned is a tool that is not being governed.


If you want a ready-made governance framework for your agency, the AI Usage Policy for recruitment agencies from Plainstart is free to download. It covers tool use, human oversight, data handling, and candidate communication in plain language you can adapt and use straight away.

This article is general guidance. It is not legal or professional advice. What applies to your agency depends on your contracts, your professional body, and whatever rules govern your practice. Your own advisers are the right people to consult on those specifics.

Free, no email required

Build your own AI usage policy in about two minutes

Answer eight questions and the full policy writes itself around your business. Copy it, download it, put it in front of staff today.

Open the policy generator