The interview scorecard: a template that survives contact with real interviewers
A scorecard exists to turn an interview into comparable evidence. Most fail because they capture a number without the reason for it — which is exactly the half you need six months later when someone asks why.
- Every rating needs a mandatory evidence field, and the evidence should be a quotation rather than a paraphrase.
- Anchors must describe observable behaviour at each scale point, not adjectives.
- Use an odd-numbered scale of five; three cannot discriminate and seven invites false precision.
- Give raters an explicit "insufficient evidence" option — forcing a rating from nothing manufactures data.
- Weight competencies before you see any candidates, and record the weighting version with the scores.
The anatomy of a scorecard that works
| Field | Why it exists | Common failure |
|---|---|---|
| Competency name | Names what is being assessed | Vague trait words that cannot be observed |
| Behavioural anchors | Makes the scale mean the same thing to two raters | Adjectives ("good", "strong") instead of described behaviour |
| Rating | The comparable output | Overall impression collected before competency ratings |
| Evidence (mandatory) | The justification, in the candidate’s words | Optional field, therefore empty; or the rater’s characterisation rather than a quotation |
| Insufficient evidence flag | Distinguishes "did badly" from "was never asked" | Absent, so raters guess and the guess becomes data |
| Rubric version | Proves which standard applied on the day | Not captured, so a later change is indistinguishable from the original |
| Rater identity and timestamp | Accountability and drift detection | Aggregated away, so systematic rater leniency is invisible |
Choosing the scale
Five points, with anchors written for 1, 3 and 5 and raters permitted to use 2 and 4 as intermediate positions. This is the range where inter-rater agreement holds up.
- Three points cannot separate a strong candidate from an adequate one, which is the separation you are usually hiring on.
- Seven or ten points invite raters to express confidence they do not have, and agreement between raters degrades sharply past five.
- Even-numbered scales force a side and can be useful where you specifically want to prevent fence-sitting, at the cost of losing a genuine middle.
- Never use a scale where the top point is defined as "exceptional" with no behavioural description — it becomes unreachable for some raters and routine for others.
A template structure to adapt
Copy the structure, replace the competencies with your own, and write anchors from your own role analysis. One block per competency.
Then a summary block, completed only after every competency block is closed:
The "factor outside the rubric" line does more work than it looks like. Hiring managers will always have considerations the rubric did not capture. Forcing them to name it converts an invisible bias into a visible, reviewable judgement — and if the named factor is unlawful, someone gets to notice before the decision lands.
Weighting, and when to set it
Set the weights before you meet any candidates and record the version. Weights adjusted after seeing the pool are indistinguishable from choosing the candidate first and the criteria afterwards, and they look exactly like that in a dispute.
- 01Rank competencies by how much variance in job performance they explain in this role, not by how much you enjoy assessing them.
- 02Give the top two or three meaningful weight and let the rest act as thresholds rather than contributors.
- 03Set explicit minimums where a competency is genuinely a floor — a candidate who cannot clear a safety-critical threshold does not get compensated by a high score elsewhere.
- 04Publish the weighting to your interviewers. Hidden weights get gamed by raters who guess at them.
What changes when it is automated
An automated scorecard has one structural advantage and one structural risk. The advantage is that the evidence field stops being optional — a scoring agent that must cite a verbatim transcript moment before it may issue a rating cannot post-hoc justify, because the evidence is retrieved before the rating exists.
The risk is false completeness. A model will produce a rating for every competency whether or not evidence exists, unless it is explicitly permitted and required to abstain. In FirstPanel each of the eight MERIT-8 agents either cites evidence or abstains, and an abstention routes a follow-up probe rather than resolving into a low score.