How do you compare candidates fairly using different tests?

7-minute read

Tadaah team meeting

You have three strong candidates and a pile of information: an interview report, a practical case study, a personality test, a reference check and perhaps a driving test or aptitude test. Everyone wants to make an “objective” choice, but in practice this actually creates confusion: scores aren’t comparable, managers assign different weights to each component, and the decision ultimately comes down to gut feeling. Fair comparison isn’t about a single perfect test, but about applying a consistent yardstick across all sources.

Why it is difficult to compare candidates fairly using different tests

Different tests measure different things, on different scales, with different margins of error. A case study shows how someone reasons in your context, an aptitude test measures general cognitive ability, and references often relate to work ethic and reliability. If you compare these results side by side without any structure, you’re comparing apples with oranges.

Group dynamics also come into play. The most persuasive interviewer, the most recent interview or a single striking detail (“that candidate had such a good story”) can subconsciously carry more weight than the evidence. Fair comparison therefore requires agreements to be made before you start, not whilst you’re evaluating.

What “fair” usually means in the selection process

Fairness does not mean treating everyone the same regardless of their role. Fairness means assessing everyone against the same relevant criteria, with the same weighting, and using the same scoring method. This makes your decision easier to explain to the team and to justify to candidates.

  • Equal opportunities: every candidate goes through the same core components.
  • A consistent yardstick: the same criteria and scoring benchmarks, regardless of who is conducting the interview.
  • Consistent interpretation: determining in advance what constitutes a “good” score.
  • Transparency: candidates know what you are assessing and why.

A practical model for making different tests comparable

Don’t start with the test, but with the role. What does a candidate need to demonstrate in the first 3 to 6 months? Translate that into a limited number of selection criteria. Only then should you choose which test or method provides the best evidence for each criterion.

Would you like to find out more about data-driven recruitment?
Discover how AI, target audience data and labour market insights can help you attract the right candidates. Download our white paper for practical insights, or feel free to get in touch for advice tailored to your organisation.

Step 1: Draw up a single set of criteria (maximum of 8) that predict success

If you have more criteria, this creates a false sense of precision and the team will still end up making decisions based on personal preference. Choose a mix of subject-specific knowledge, work behaviour and contextual factors. For blue-collar, technical, logistics or maritime roles, criteria such as safety, accuracy and problem-solving are often more important than presentation.

  • Technical skills (e.g. fault diagnosis, welding, order picking, key account management)
  • Work behaviour (safety, quality, ownership, pace)
  • Collaboration (handover, coordination, processing feedback)
  • Learnability (the speed at which a person improves following instruction)
  • Contextfit (shift work, travel distance, physical strain, rules)

Step 2: Assign a primary and secondary measurement source to each criterion

Every source has its limitations. An interview is a strong tool for assessing motivation and behaviour, but less so for actual performance. A case study is a strong tool for assessing thought processes, but less so for real-world collaboration. By selecting a primary source and a backup source for each criterion, you can prevent any single aspect from being the sole determining factor.

Introduce the table as a decision-making tool: this explains why you use multiple tests and how they combine to form a single assessment.

Criterion Primary measurement Secondary measurement What do you record as evidence?
Problem-solving skills Case study Structured interview (STAR) Step-by-step plan, hypotheses, checks, focus on results
Safety awareness Case studies + follow-up enquiries regarding incidents Qualifications/work experience Identifying risks, measures, awareness of standards
Accuracy/quality Work sample or debugging task Reference check Checkpoints, error prevention, coping with pressure
Collaboration Behavioural interview Reference check Examples of coordination, conflict and feedback
Learnability Mini-assignment with a feedback session Interview question about learning Reflection, adjustment, impact following feedback

Step 3: Normalise scores to a single scale (e.g. 1–5)

This is the key to comparing candidates fairly across different tests. You don’t need to make everything statistically perfect, but you do need to be consistent. Convert each test to the same final score scale and document how you do this.

  • Case study: rubric with 1–5 criteria (ranging from “overlooks key risks” to “works in a structured manner and checks assumptions”).
  • Interview: score for each criterion based on evidence, not on general impression.
  • Personality questionnaire: do not interpret the results as “right or wrong”, but rather as “risks and areas for attention”, using a fixed decision rule.
  • References: standard questions and the same scoring criteria (e.g. reliability, cooperation, pace).

Avoid adding together raw test scores from different tests. If you do use test percentiles, create a single conversion table and always use that one.

Step 4: Use weighting, but keep it simple

Weighting is useful when not everything is equally important. Too much maths means teams will abandon the model and fall back on “whatever feels right”. You should therefore choose a maximum of three key factors and give the rest equal weighting.

  • For operators/technicians: safety and problem-solving take precedence over communication.
  • In logistics: accuracy and speed take precedence over creativity.
  • In sales: needs analysis and discipline are more important than charm.

How to minimise bias when combining sources

Even with a good scorecard, unconscious bias can still creep in. The solution often lies in establishing clear processes: when do you award marks, who sees what information and when, and how do you discuss discrepancies? This is what we call ‘interview hygiene’, but applied to the entire selection process.

Let assessors mark the work independently first

Ask everyone to note down a score and supporting evidence for each criterion before you meet. This will prevent a single dominant opinion from dictating the discussion. It also makes it possible to discuss differences without anyone “losing out”.

For each criterion, discuss the evidence, not the person

A useful question to ask in evaluations is: “What specific evidence did you see that explains your score?” This brings the discussion back to observation. Statements such as “I just don’t see it” rarely result in a fair comparison.

Treat tests as input, not as a final verdict

Personality and aptitude tests, in particular, are often used in too absolute a manner. Yet many of these tools are intended to aid the interview process: where might potential risks lie, how does a person cope with pressure, how do they learn? If you use them as a ‘knock-out’ criterion, explain what role that profile plays in this role and why.

Make knock-outs explicit in advance

A fair comparison becomes skewed if you introduce new requirements during the recruitment process. You should therefore establish in advance what is non-negotiable. Examples include a safety certificate, a driving licence, compulsory shift work or a minimum score on safety awareness.

What to do and what not to do with references and test data

References and test results are often the least well-organised, yet they have a significant influence in the final stage. You should therefore keep them brief and standardised. Two targeted reference questions with fixed scoring criteria are more valuable than a long conversation that mainly confirms what you already thought.

Standardise your reference check

  • Use the same questions for each candidate applying for the same role.
  • Ask for specific examples: “When did things go wrong, and what did the candidate do then?”
  • First, ask the referee to outline the context (team size, type of work, workload).

Please be aware of privacy and retention periods

You collect application data, test results and notes. This falls under data protection legislation and requires a clear legal basis, security measures and retention periods. The Dutch Data Protection Authority provides practical guidance on how to handle application data, including retention and informing candidates; see Explanation from the Dutch Data Protection Authority regarding job application data.

A brief checklist for consistent decision-making

If you want to make a fairer comparison as early as tomorrow, these are the quickest steps you can take that will make the biggest difference. They take very little time, but will help eliminate a lot of the noise from your decision.

  • A maximum of 6–8 criteria, linked to success in the role.
  • For each criterion: a primary measurement method and one backup source.
  • Everything on a single scale (1–5) with clear scoring benchmarks.
  • Interviewers mark independently and share evidence.
  • Set the knock-outs in advance (and do not change them).
  • Test results: always interpret them in context, not as absolute truths.

Would you like to review this process for a specific target group (blue-collar, sales or maritime) and translate it into scorecards, case studies and a candidate journey with fewer dropouts? In that case, a data-driven approach such as the one used by FosFor often works well. Please feel free to get in touch so that we can work together to build a selection process that compares candidates more fairly and leads to a consensus-based decision more quickly.

We would like to get in touch

Get in touch

Good people don’t look for job vacancies. Download the white paper and find out how you can reach them anyway.