AI interviews versus psychometric testing: what each actually predicts
Psychometric testing and structured interviewing are frequently presented as competing answers to the same question. They are not — they measure different things, carry different legal exposure, and the sensible posture for most Australian employers is to understand which risk they are taking on.
- Cognitive ability tests predict well across roles and carry the best-documented adverse impact of any common selection method.
- Personality inventories are cheap and popular, and their standalone predictive validity for job performance is modest.
- Structured interviews predict well and produce something the other two do not: an explainable individual record.
- In Australia, personality and cognitive instruments raise a specific disability discrimination exposure via inferred traits and timed conditions.
- If you combine methods, weight them explicitly in advance and monitor adverse impact at each stage rather than only at the end.
What each actually measures
| Method | Measures | Predictive strength | Principal risk |
|---|---|---|---|
| Cognitive ability test | General mental ability, reasoning speed | Consistently among the strongest single predictors across role types | The best-documented adverse impact of any common method, and timed conditions raise disability adjustment obligations |
| Personality inventory | Self-reported trait tendencies | Modest on its own; conscientiousness is the most consistently useful facet | Self-report is coachable and fakeable; inferred traits invite a disability argument |
| Structured interview | Job-relevant behavioural evidence against anchored criteria | Strong, and among the best-evidenced improvements available to a hiring process | Costly in assessor hours, which is why structure collapses under volume |
| Work sample | Demonstrated capability on representative work | Strong where the job is definable as a task | Completion effort, and simulation fidelity for judgement-heavy roles |
The Australian legal exposure is different for each
This is the part that rarely appears in vendor comparisons and that Australian employment counsel raise first.
Inferred traits and disability
A personality inventory that infers emotional stability, resilience or stress tolerance is producing an assessment that overlaps with mental health. Screening a candidate out on that inference is difficult to distinguish from screening them out on a disability, and the Disability Discrimination Act 1992 and s.351 of the Fair Work Act both apply. The instrument may be well validated; the question in a dispute is whether the inferred trait was job-related and whether the attribute formed part of the reasons.
Timed conditions and adjustments
Timed cognitive testing creates an adjustment obligation. Candidates with a range of disabilities are entitled to reasonable adjustments, which for a speeded test means extended time — and extended time changes what a speeded test measures. Plan for how you compare an adjusted result to an unadjusted one before you need to, not afterwards.
Explainability
The decisive difference in Australia. Under s.361 of the Fair Work Act you must prove a protected attribute formed no part of the reasons. A percentile score on a normed instrument tells a tribunal where a candidate sat against a norm group; it does not tell them what the candidate did or said that justified the decision. A structured interview with verbatim evidence per rating does.
Candidate acceptance
Face validity — whether the assessment looks related to the job — drives both completion and post-rejection goodwill, and it varies sharply between these methods.
- Work samples have the highest acceptance: candidates accept being asked to do the job.
- Structured interviews are well accepted where the questions are visibly job-related.
- Cognitive tests are tolerated but resented at volume, particularly when a candidate cannot see the connection to a frontline role.
- Personality inventories attract the most scepticism, especially when candidates suspect the "right" answers, which they usually do.
In consumer-facing industries the acceptance question is a brand question. Your applicants shop with you, and an assessment that felt arbitrary is remembered as an interaction with the company, not with a vendor.
Combining them defensibly
- 01Decide what each instrument is for and write it down. "We use a cognitive test because reasoning under time pressure is a genuine requirement of this role" is a defensible sentence. "We use it because everyone does" is not.
- 02Weight in advance and version the weighting. Weights adjusted after seeing the pool look exactly like choosing the candidate first.
- 03Monitor adverse impact at each stage transition, not just end to end. Cognitive testing’s impact typically appears at its own stage and can be masked in the aggregate.
- 04Build the adjustment path before the first candidate needs it, and record adjustments in the decision file as evidence of compliance.
- 05Keep one instrument in the stack that produces individual, quotable evidence. It is what you defend with, and neither a percentile nor a trait profile provides it.