HiringRecruitingInterviews

Interview Scorecard: Template, Examples, and the Six Parts That Make One Work (2026)

The Rankid Team·August 30, 2026·16 min read
Two interviewers reviewing the same candidate reach opposite verdicts, a strong yes based on great communication and a no based on seeming vague, next to a scorecard that resolves the disagreement by rating four weighted criteria from 1 to 4 against written anchors with evidence recorded for each

Two interviewers spend 45 minutes each with the same candidate. One writes “strong yes, excellent communicator, would hire.” The other writes “no, vague, could not get a straight answer.” Neither of them is lying, and neither is careless. They reached opposite conclusions because nobody had written down what the interview was supposed to measure, so each of them measured something reasonable and different. The document that prevents this has a name, and most teams build it at the wrong end of the process.

Quick answer

An interview scorecard is four to six weighted criteria, each with written descriptions of what a 1 and a 4 look like, plus a box for the evidence behind every rating. Build it before you write the questions, use 1 to 4 with anchors rather than 1 to 5 with adjectives, and have everyone submit independently before the debrief opens. The part almost everyone skips: the criteria should be the same ones you screened resumes against, so the screen and the interview are one rubric instead of two opinions. Score your applicant pool against role criteria with Rankid. First 5 resumes free, no signup.

What is an interview scorecard?

An interview scorecard is a short document, agreed before anyone is interviewed, that states the criteria a candidate is being evaluated against, how much each one counts, and what different levels of performance against it look like in observable terms. Interviewers rate the candidate on each criterion and record what they actually heard that justifies the rating.

You will see the same artefact called a candidate scorecard, a hiring scorecard, a recruiting scorecard, an interview evaluation form or an interview scoring sheet. The names are interchangeable. What is not interchangeable is whether the criteria are specific to the role and written down in advance, because that is the only feature doing any work.

Three things a scorecard is not:

  • It is not a record of the decision. A form completed after you already know your answer is documentation of a conclusion, not a method for reaching one.
  • It is not a generic competency sheet. Communication, attitude and experience rated one to five, reused for every role from warehouse operative to CFO, measures whichever of those things happened to stand out.
  • It is not a substitute for judgement. It constrains what your judgement is allowed to be about, then makes the resulting judgements comparable across interviewers.

Why two interviewers disagree about the same candidate

Two interviewers reviewing the same candidate reach opposite verdicts, a strong yes based on great communication and a no based on seeming vague, next to a scorecard that resolves the disagreement by rating four weighted criteria from 1 to 4 against written anchors with evidence recorded for each

The disagreement at the top of this article is not a people problem. It is what happens when an evaluation has no defined target. Left to itself, an unstructured interview tends to converge on how confident, articulate and familiar a candidate seemed, because those are the qualities that are legible in a conversation whether or not the job depends on them. The consistent finding across decades of selection research is that structured interviews, meaning the same questions rated against defined criteria, predict job performance considerably better than unstructured ones. The scorecard is what makes an interview structured on the rating side, the same way a fixed question set does on the asking side.

There is a subtler failure that scorecards fix, and it survives even in teams that think they are already structured. Two interviewers both give a candidate a 3. That looks like agreement, so nobody examines it. One of them meant “did the job adequately at a smaller company and would need support here”, the other meant “nothing went wrong in the conversation”. Without written anchors, the number is a shared symbol for two unshared judgements, and the average of them means nothing at all.

Disagreement is the point, not the problem

A good scorecard produces more visible disagreement, not less, because it turns vague consensus into specific conflict you can resolve by pointing at evidence. A hiring process where everyone always agrees is usually one where the first confident opinion in the room is doing all the work.

The six parts of a scorecard that actually works

Most scorecard templates supply two of these six and wonder why the output is thin. Each criterion on the sheet needs all six.

A single interview scorecard criterion broken into its six required parts: the criterion name, one line on why it predicts performance in this job, the weight as a percentage of the decision, written anchors describing what a score of 1 and a score of 4 look like in observable behaviour, an evidence box recorded before the rating, and the rating itself submitted independently before the debrief
  • 1. The criterion.Something the job genuinely requires, phrased narrowly enough to be observed. “Debugging an unfamiliar system” is a criterion. “Technical ability” is a category.
  • 2. Why it predicts. One line tying it to the actual work. If you cannot write that line, the criterion is inherited rather than chosen, and it is diluting the weights of the ones that matter.
  • 3. The weight. A percentage of the decision, set at kickoff, summing to 100 across the sheet.
  • 4. The anchors. A written description of what a 1 looks like and what a 4 looks like, in behaviour you could observe rather than qualities you could attribute.
  • 5. The evidence box. What the candidate actually said or did, written before the rating is chosen. This ordering matters more than it sounds.
  • 6. The rating. Submitted independently, before the debrief opens.

Part five carries surprising weight. Asking interviewers to write the evidence first, then pick the number, forces the rating to be a conclusion from something specific. Reversing the order lets people pick a number from impression and then hunt for a justification, which is the same process with the audit trail attached.

A free interview scorecard template you can copy

Copy this into a doc, a spreadsheet or your ATS. It takes about twenty minutes per role to fill in, once, and every interview for that role reuses it.

1

Header block

Role and requisition ID, candidate name, interviewer name, interview stage, date, and the criteria this interviewer is responsible for. That last field matters on panels: assign each criterion to one or two interviewers rather than having everyone rate everything, which produces four shallow opinions instead of two evidenced ones.

2

Criteria block, repeated four to six times

Criterion: ____. Why it predicts performance here: ____. Weight: ___%. A 1 looks like: ____. A 4 looks like: ____. Evidence heard: ____. Rating (1 to 4): ___.

3

Knockouts block

A short list of pass or fail requirements that are not weighted and not tradeable: the licence, the legal right to work, the mandatory certification, the shift availability. Keeping these out of the weighted section prevents the most damaging scorecard error, which is a strong average quietly outvoting a hard requirement.

4

Recommendation block

One overall call on a four-point scale with no midpoint: strong no, no, yes, strong yes. Then two free-text lines that are worth more than the rest of the form combined: “The strongest evidence for hiring is…” and “What would change my mind is…”. The second line is what makes a debrief productive rather than positional.

Weighted total, and when to ignore it

Multiply each rating by its weight and sum for a single comparable figure out of 4. Treat it as a sorting aid, never as the decision. A 3.4 built on a 1 in the highest-weighted criterion is not a better candidate than a 3.1 with no weak spot, and any process that cannot see that has replaced judgement with arithmetic.

Interview scorecard examples with real anchors

Anchors are the part people struggle to write, so here are worked examples from two very different roles. Notice that every anchor describes something a candidate does, not something they are.

Customer support specialist, criterion: de-escalating an angry customer (weight 30%)

  • 1: Describes the customer as the problem, or the example given is one where a policy was simply restated until the customer gave up.
  • 4: Gives a specific incident, states what the customer actually wanted underneath the complaint, describes what they conceded and what they held, and can say what the outcome was including whether the customer stayed.

Account executive, criterion: qualifying out of a bad deal (weight 25%)

  • 1: Cannot recall walking away from a deal, or frames qualification purely as budget confirmation.
  • 4: Names a specific deal they killed while it was still live, the signal that prompted it, roughly what it was worth, and what they did with the time instead.

The test for an anchor is whether two people reading it and hearing the same answer would land on the same number. “Communicates well” fails that test instantly. “Can state what the customer wanted underneath the complaint” passes it. If you want the question side of this, interview questions to ask candidates covers which questions produce evidence against criteria like these, and the screening interview guide covers the shorter version for a first call.

Where the criteria should come from, and why it is not the interview

Here is the structural mistake, and it is nearly universal. Teams build the scorecard when they schedule the first onsite. By then, hundreds of applicants have already been sorted into a shortlist using criteria that were never written down, and the scorecard is being asked to evaluate a group that was selected on a different basis than it is about to be measured on.

A comparison of two hiring processes: in the broken version the resume screen, the phone screen and the onsite each invent their own criteria so the stages do not connect, while in the working version one set of five weighted role criteria runs through all three stages, with the resume screen producing a 0 to 100 match score, the phone screen targeting the criteria the resume could not evidence, and the onsite rating the same criteria from 1 to 4

The criteria already exist. When somebody reads a resume and decides it is worth a call, they are applying criteria, just privately and inconsistently. Writing the scorecard at the point where you open the requisition, rather than at the point where you start interviewing, means one set of criteria governs the entire process. That single change produces three effects:

  • The interview stops re-litigating the screen. If the resume already evidenced eight years of the required stack, the interview should not spend twenty minutes confirming it. Interview for what paper cannot establish.
  • The screen tells you what to ask. A candidate who scores well overall but has no evidence at all against criterion three has just written your question list for you. The gaps are the agenda.
  • Your funnel becomes measurable. One rubric across stages means selection rates per criterion, per stage, which is the raw material for both recruitment metrics and the fairness checks in adverse impact analysis.

The practical blocker is capacity, not conviction. Applying five weighted criteria to 400 applicants by hand is not something a recruiter carrying six requisitions can do, which is exactly why the shortlist ends up built by skimming the first 40 and the criteria only get written down later. That is the constraint worth removing first, and it is the same argument made in how to screen resumes in bulk and resume screening criteria.

Start the scorecard at the resume, not at the interview

Upload your full applicant batch and paste the job description. Rankid scores every candidate 0 to 100 against the role's criteria and shows the skills and keywords each one matches and misses, so your shortlist is built on the same criteria your interview scorecard rates, and every interview starts with a list of exactly what to probe. Up to 200 resumes per batch, first 5 free, no signup.

Score your applicant pool free

Rating scales: why 1 to 4 beats 1 to 5

Use an even-numbered scale. The midpoint of a five-point scale is where undecided interviewers park candidates they did not really evaluate, and a sheet full of 3s is indistinguishable from no evaluation at all. Removing it forces a direction, which is the smallest possible amount of commitment to ask of someone who just spent an hour with a person.

  • 1 Clearly below the bar for this role. You saw evidence against.
  • 2 Below the bar, but arguable, or the evidence was thin.
  • 3 Clearly meets the bar. Would be fine doing this part of the job.
  • 4 Exceeds it, with a specific thing you can point at.

Two rules that matter more than the scale itself. First, allow “no evidence gathered” as a distinct entry, separate from a 2. An interviewer who never got to a criterion must be able to say so, because a guessed 2 is worse than a blank and will be averaged in as though it were a finding. Second, never let an interviewer see anyone else's rating before submitting their own.

Set the weights before you meet anyone

Weights assigned after interviews have started are rationalisation. Everyone can construct a weighting under which their preferred candidate wins, and everyone does it sincerely. Agree the percentages at the kickoff meeting with the hiring manager, in the same session where you agree the job description, and write them down before a single resume is read.

The kickoff question that produces honest weights is not “what matters for this role?”, because the answer is always everything. Ask instead: if you had to hire someone who was a 2 on exactly one of these, which one would it be? That surfaces the real ranking in about ninety seconds, and it usually reveals that a criterion everyone listed as essential is one the team would happily trade. A related discipline, deciding which requirements are genuinely required at all, is the substance of skills-based hiring.

Running a debrief that the scorecard survives

A scorecard filled in honestly and then discussed badly produces the same outcome as no scorecard. The debrief is where the method is usually lost.

1

Collect before you convene

Every scorecard submitted before the meeting starts, with no visibility of others. If someone has not submitted, the debrief does not begin. This single rule does more than the rest of the process combined, because it prevents the first confident voice from anchoring the room.

2

Read the spread, not the average

Put the ratings up per criterion. Where everyone agrees, move on quickly. The value is entirely in the criteria where a 1 sits next to a 4, because one of those interviewers heard something the other did not.

3

Resolve disagreement with evidence, in reverse seniority

Ask the two divergent raters to read out their evidence box, not their conclusion. Most gaps close immediately, because one person asked a follow-up the other did not. Have the most junior person speak first throughout.

4

Decide, then record what would have changed it

Write one line on the deciding factor. Six months later, when you are trying to work out whether your criteria predict anything, that line is the only data you will have. It is also the beginning of a quality-of-hire loop rather than a filing exercise.

Spreadsheet, ATS, or screening tool

The format matters less than the enforcement, but the trade-offs are real:

  • Spreadsheet.One tab per candidate, criteria as rows, interviewers as columns, weighted total computed automatically. Free, flexible, and the version most teams should start with. The weakness is that nothing stops someone opening a colleague's column before submitting.
  • ATS scorecards. Most modern systems, including Greenhouse-style tools, support per-stage scorecards with required fields and locked submissions, which solves the independence problem properly. The risk is the default template: generic competencies shipped by the vendor, reused across every role, which is the failure mode from the top of this article wearing better software. Replace them per role. Options in our applicant tracking systems list.
  • Screening tool for the paper stage. The resume half of the rubric is the part that does not scale by hand, so this is where automation earns its place, with the interview half staying human. The boundary between the two is drawn in recruitment automation.

Seven ways interview scorecards fail

  • Scorecard theatre. Completed after the decision, to document it. Detectable by a suspicious absence of low scores for hired candidates.
  • Twelve criteria. Weights become noise, and interviewers rate things they gathered no evidence on rather than leave a blank.
  • Adjective anchors.“Excellent communicator” is a rating pretending to be a definition. Anchors must describe observable behaviour.
  • Averaging over a knockout. A 3.6 weighted total on a candidate who cannot legally do the job. Keep hard requirements out of the weighted section entirely.
  • Resume halo. Interviewers who read the resume closely arrive with a prior and score the conversation to match. Distribute criteria across the panel so nobody rates everything.
  • Rating the interview, not the candidate. Punishing nerves, an accent or an unpolished answer rewards interview practice, which is not on the scorecard and rarely on the job.
  • No connection to the screen. The most expensive one. Rigorous scoring applied to a shortlist that was assembled by skimming means you have measured the wrong forty people precisely, which is the coverage problem described in how to shortlist candidates from resumes.

Do scorecards make hiring fairer, honestly?

Partly, and it is worth being precise about the mechanism rather than claiming more than is true. A scorecard narrows what an interviewer is permitted to consider and requires each rating to be tied to stated evidence, which shrinks the space in which impressions about confidence, background or similarity operate. It does not make anyone neutral. A scorecard whose criteria include “culture fit” with no anchors has done nothing at all except give a gut feeling a number and a paper trail.

The second benefit is more reliable than the first and gets discussed less. Consistent criteria make the process measurable. When every candidate has been rated on the same five things, you can compute selection rates by stage and by group and see where a stage is filtering unevenly, using the four-fifths method in adverse impact. Free-text interview notes cannot be analysed at all, so a process without scorecards is not neutral, it is merely unexamined. It is also the difference between a defensible hiring decision and one that rests on the recollection of whoever was in the room.

Key takeaways

  • An interview scorecard is four to six weighted, role-specific criteria with written anchors and an evidence box, agreed before anyone is interviewed.
  • Write the evidence first and the rating second, or the number becomes an impression looking for a justification.
  • Use 1 to 4, not 1 to 5. The midpoint is where unevaluated candidates get parked.
  • Anchors must describe observable behaviour. If two readers would score the same answer differently, the anchor is not written yet.
  • Set weights at kickoff by asking which single criterion the hiring manager would accept a 2 on.
  • Keep hard requirements as knockouts outside the weighted total, so a strong average cannot outvote them.
  • Build the scorecard when you open the req, not when you schedule the onsite, so the resume screen and the interview rate the same criteria.
  • Rankid scores up to 200 resumes per batch against one job description. First 5 free, no signup.

Bottom line: an interview scorecard is not a form, it is an agreement about what is being measured, and its value is set almost entirely by when you write it. Written the week of the onsite, it makes a late-stage conversation tidier. Written the day the requisition opens, it decides who enters the funnel, what each interview is for, and whether the numbers coming out the other end mean anything. Same document, completely different instrument. Score your next applicant batch against your role criteria with Rankid and let the scorecard start where the candidates do.

Frequently asked questions

What is an interview scorecard?

An interview scorecard is a short written document, agreed before anyone is interviewed, that lists the four to six criteria which genuinely predict success in a role, assigns each a weight, and describes in observable terms what a low and a high score look like. Interviewers rate the candidate against those criteria and record the evidence behind each rating. It goes by several names, including candidate scorecard, hiring scorecard, interview evaluation form and interview scoring sheet, and they all describe the same artefact. Its purpose is not documentation. It is to make two interviewers' opinions comparable, which they are not when each person is silently measuring something different.

What should be included on an interview scorecard?

Six things per criterion: the criterion itself, one line on why it predicts performance in this specific job, a weight expressed as a percentage of the decision, written anchors describing what a 1 and a 4 look like in observable behaviour, a free-text box for the evidence the interviewer actually heard, and the rating. The scorecard as a whole also needs identifying fields for role, candidate, interviewer and stage, plus a single overall recommendation on a four-point scale with no middle option. Keep it to four to six criteria. Scorecards with twelve criteria get filled in carelessly, which is worse than not having one.

What is the difference between an interview scorecard and an interview evaluation form?

In practice they are the same document under different names, but the names carry different habits. An evaluation form is often a generic sheet used for every role, with broad headings such as communication, attitude and experience rated one to five. A scorecard is built per role from the criteria that actually predict success in that job, with weights and written anchors. The distinction that matters is not the label but whether the criteria are role-specific and defined before anyone is interviewed. A generic form filled in after the decision has already been made is paperwork, whatever it is called.

What rating scale should an interview scorecard use?

Use 1 to 4 with written anchors, not 1 to 5. An even-numbered scale removes the midpoint, and the midpoint is where undecided interviewers park candidates they have not really evaluated. A useful set of definitions is 1 for clearly below the bar, 2 for below but arguable, 3 for clearly meets the bar, and 4 for exceeds it in a way you can point at. The scale matters far less than the anchors. Two interviewers who both write 3 without anchors frequently mean different things, so the number gives you false agreement rather than real comparison.

How many criteria should be on a hiring scorecard?

Four to six. Below four you are usually missing something the job actually requires; above six the weights become meaningless and interviewers start rating criteria they gathered no evidence on, which is the most common way scorecards degrade into fiction. If a role seems to need ten criteria, most of them are either restatements of the same underlying ability or things you can check on paper rather than in an interview. Move the checkable ones to the resume screen and keep the interview for what only a conversation can establish.

Do interview scorecards reduce bias?

They reduce it rather than remove it, and the mechanism is worth understanding. Scorecards constrain what interviewers are allowed to consider and force each rating to be tied to stated evidence, which narrows the space in which impressions about confidence, accent, background or similarity can operate. They do not make an interviewer neutral, and a scorecard with vague anchors such as good culture fit simply relabels a gut feeling as a score. The second benefit is measurement: once every candidate is rated against the same criteria, you can calculate selection rates by stage and detect adverse impact, which is impossible when every evaluation is a paragraph of free text.

Can you use a scorecard to screen resumes as well as interviews?

You should, and it is the step most teams miss. The criteria that decide who gets hired should be the same criteria that decide who gets a call, otherwise the interview either re-litigates the screen or measures something unrelated to it. In practice the resume screen produces a score against role criteria and the interview produces a rating against the same criteria, giving you one continuous rubric rather than two disconnected judgements. The screen also tells you what to interview for: the criteria a candidate's resume could not evidence are exactly the ones the conversation should target.

How do you stop interviewers filling in the scorecard after they have decided?

Require independent submission before the debrief begins, and make the debrief read submitted scores rather than collect them. This is a process control, not a trust issue. If interviewers discuss the candidate first, the first confident opinion in the room anchors everyone else, and the scorecards written afterwards will agree with it, which looks like consensus and is actually contamination. Two further habits help: ask each interviewer to record the evidence before the rating, and ask the most junior interviewer to speak first in the debrief.

Written by the The Rankid Team. See more in our blog, or check your resume against a job now.