Adverse impact testing: the four-fifths rule, done properly
The four-fifths rule is easy to implement almost right, and every "almost" produces a confident, wrong legal conclusion. Here is the calculation, the two mistakes that matter, and what the result actually tells you.
- Selection rate is selected divided by applicants, computed per group.
- The impact ratio compares each group’s selection rate to the group with the highest rate — not to the majority group.
- A ratio below 0.80 is generally regarded as evidence of adverse impact; above it is not a safe harbour.
- Small groups produce wild ratios. Report the number, mark it unreliable, and do not act on it as a finding.
- Monitor per requisition and before shortlists ship, not quarterly after the harm.
The calculation
The four-fifths rule comes from the US Uniform Guidelines on Employee Selection Procedures at 29 CFR 1607.4(D). It has no direct statutory force in Australia, but it is the analysis regulators, plaintiffs and auditors everywhere reach for, and it is the methodology underlying bias-audit regimes such as New York City’s Local Law 144.
The worked example from the Guidelines themselves
| Group | Applicants | Selected | Selection rate | Impact ratio |
|---|---|---|---|---|
| Men | 100 | 80 | 80% | 1.00 (reference) |
| Women | 50 | 20 | 40% | 0.50 — below threshold |
The reference-group mistake that hides real impact
This is the error that quietly makes a discriminatory process look clean, and it is extremely common in home-built dashboards.
The reference group is the group with the highest selection rate. It is not the majority group, and it is not the group you consider the default. Anchoring to the majority produces ratios above 1.00 for everyone who outperforms it and conceals the group that is actually being screened out.
| Group | Applicants | Selected | Rate | Ratio vs majority (wrong) | Ratio vs highest (correct) |
|---|---|---|---|---|---|
| A (majority) | 300 | 90 | 30% | 1.00 | 0.50 — flagged |
| B | 50 | 30 | 60% | 2.00 | 1.00 (reference) |
| C | 50 | 8 | 16% | 0.53 | 0.27 — flagged |
Small groups, and the number you must not act on
With seven applicants in a group, one hiring decision swings the ratio by tenths. A dashboard that reports "adverse impact detected" from three candidates is not measuring fairness; it is measuring noise, and it will train your team to ignore the alerts that matter.
- Below a minimum group size, report the number but mark it explicitly unreliable, and never let it trigger a finding.
- Aggregate across requisitions of the same role family to reach usable sample sizes, but keep the per-requisition view for the alerting that has to happen before a shortlist ships.
- Be careful aggregating across genuinely different roles — a combined ratio can hide a serious problem in one requisition behind clean numbers in others.
- Report the counts alongside every ratio. A ratio without an n is an invitation to over-read it.
What a passing ratio does not mean
Three things the four-fifths rule will not tell you, all of which get claimed anyway.
- 01It does not tell you that you are compliant. It is a rule of thumb for triggering scrutiny, not a safe harbour. The Guidelines themselves say smaller differences may still constitute adverse impact where they are statistically significant.
- 02It does not tell you the cause. A ratio below threshold says the outcome differs by group; it says nothing about which stage, criterion or model behaviour produced the difference. Locating that requires stage-by-stage analysis.
- 03It does not discharge an Australian legal burden. Under s.361 of the Fair Work Act the question is whether a protected attribute formed part of the reasons for an individual decision. An aggregate ratio is context, not evidence about that candidate.
Operationalising it
- 01Define the stages. Application, screen, interview, shortlist, offer. Compute rates at each transition — impact often appears at one stage and washes out in the aggregate.
- 02Decide your attributes and your data source. Only what candidates voluntarily provide, held separately from assessment data and never visible to the assessment.
- 03Compute per requisition, continuously, and alert before shortlists ship rather than in a quarterly report.
- 04Record every alert and its resolution. The fact that you monitored, and what you did, is itself part of the defence.
- 05When a threshold is crossed, investigate the criterion rather than the candidate. The question is which requirement produced the difference and whether it is genuinely job-related.
- 06Re-test after any rubric or model change. A change that improves average quality can degrade a group’s outcomes.