HomeGuidesHow to build an interview scorecard
Guides · Interviewing

The interview scorecard: a template that survives contact with real interviewers

A scorecard exists to turn an interview into comparable evidence. Most fail because they capture a number without the reason for it — which is exactly the half you need six months later when someone asks why.

Updated 23 August 20269 min readFirstPanel research team
Key takeaways
  • Every rating needs a mandatory evidence field, and the evidence should be a quotation rather than a paraphrase.
  • Anchors must describe observable behaviour at each scale point, not adjectives.
  • Use an odd-numbered scale of five; three cannot discriminate and seven invites false precision.
  • Give raters an explicit "insufficient evidence" option — forcing a rating from nothing manufactures data.
  • Weight competencies before you see any candidates, and record the weighting version with the scores.

The anatomy of a scorecard that works

FieldWhy it existsCommon failure
Competency nameNames what is being assessedVague trait words that cannot be observed
Behavioural anchorsMakes the scale mean the same thing to two ratersAdjectives ("good", "strong") instead of described behaviour
RatingThe comparable outputOverall impression collected before competency ratings
Evidence (mandatory)The justification, in the candidate’s wordsOptional field, therefore empty; or the rater’s characterisation rather than a quotation
Insufficient evidence flagDistinguishes "did badly" from "was never asked"Absent, so raters guess and the guess becomes data
Rubric versionProves which standard applied on the dayNot captured, so a later change is indistinguishable from the original
Rater identity and timestampAccountability and drift detectionAggregated away, so systematic rater leniency is invisible

Choosing the scale

Five points, with anchors written for 1, 3 and 5 and raters permitted to use 2 and 4 as intermediate positions. This is the range where inter-rater agreement holds up.

  • Three points cannot separate a strong candidate from an adequate one, which is the separation you are usually hiring on.
  • Seven or ten points invite raters to express confidence they do not have, and agreement between raters degrades sharply past five.
  • Even-numbered scales force a side and can be useful where you specifically want to prevent fence-sitting, at the cost of losing a genuine middle.
  • Never use a scale where the top point is defined as "exceptional" with no behavioural description — it becomes unreachable for some raters and routine for others.

A template structure to adapt

Copy the structure, replace the competencies with your own, and write anchors from your own role analysis. One block per competency.

Then a summary block, completed only after every competency block is closed:

The "factor outside the rubric" line does more work than it looks like. Hiring managers will always have considerations the rubric did not capture. Forcing them to name it converts an invisible bias into a visible, reviewable judgement — and if the named factor is unlawful, someone gets to notice before the decision lands.

Weighting, and when to set it

Set the weights before you meet any candidates and record the version. Weights adjusted after seeing the pool are indistinguishable from choosing the candidate first and the criteria afterwards, and they look exactly like that in a dispute.

  1. 01Rank competencies by how much variance in job performance they explain in this role, not by how much you enjoy assessing them.
  2. 02Give the top two or three meaningful weight and let the rest act as thresholds rather than contributors.
  3. 03Set explicit minimums where a competency is genuinely a floor — a candidate who cannot clear a safety-critical threshold does not get compensated by a high score elsewhere.
  4. 04Publish the weighting to your interviewers. Hidden weights get gamed by raters who guess at them.

What changes when it is automated

An automated scorecard has one structural advantage and one structural risk. The advantage is that the evidence field stops being optional — a scoring agent that must cite a verbatim transcript moment before it may issue a rating cannot post-hoc justify, because the evidence is retrieved before the rating exists.

The risk is false completeness. A model will produce a rating for every competency whether or not evidence exists, unless it is explicitly permitted and required to abstain. In FirstPanel each of the eight MERIT-8 agents either cites evidence or abstains, and an abstention routes a follow-up probe rather than resolving into a low score.

FAQ

Frequently asked

What should an interview scorecard include?+

A named competency with an observable definition, behavioural anchors at each scale point, the rating, a mandatory evidence field containing a verbatim quotation, an explicit insufficient-evidence option, the rubric version, and the rater identity and timestamp.

What rating scale should I use?+

Five points with anchors written for 1, 3 and 5. Three points cannot discriminate usefully; seven or more degrades agreement between raters without adding real information.

Should interviewers see each other’s scores?+

Not before submitting their own. Independent scoring first, then discussion — otherwise the first score anchors the rest and you have collected one opinion recorded several times.

How long should we keep scorecards?+

In Australia, set a six-year floor for the decision record. A rejected applicant is not caught by the 21-day dismissal clock, so the general six-year limitation period applies to a general protections claim about the decision.

See it on one of your own roles

Pick your longest-open requisition. First interviewed shortlist in about two weeks — keep the reports either way.