9 min readMarcus Thorne

Updated on

Fixing the scoring gap in structured interviews

How specific behavioral rubrics prevent hiring teams from reverting to gut feelings after the candidate leaves the room.

Fixing the scoring gap in structured interviews

{ "body_markdown": "## The failure of the silent rubric\n\nMost recruiting teams believe they are conducting structured interviews because they use a standard list of questions. They gather in a conference room or a Zoom call, ask the same five behavioral questions to every candidate, and take diligent notes. However, the structure usually ends the moment the candidate stops talking. \n\nWhen the hiring team meets for the debrief, the conversation often dissolves into subjective statements like "they felt like a good culture fit" or "I am not sure they have the right energy for this team." These phrases are signals that your structured interview process has a scoring gap. Without a pre-defined scale for what a good answer actually looks like, your team is simply using a structured format to justify an unstructured decision.\n\nThis gap creates two major risks. In the United States, it opens the door to litigation if a candidate can prove that hiring decisions were based on inconsistent criteria. In Europe, especially under GDPR and local labor laws in countries like Germany, candidates have a right to understand the basis of a rejection. If your only record is a set of vague notes, you lack a defensible audit trail. More importantly, you end up hiring the best talker rather than the best performer.\n\n## Moving from questions to evidence\n\nTo fix this, you must stop focusing on the questions and start focusing on the evidence. A true structured interview requires a three-part framework for every competency you test: the question, the positive indicators, and the negative indicators. \n\nTake the competency of "conflict resolution" for a Project Manager role. A typical question might be: "Tell me about a time you had a disagreement with a stakeholder about a project timeline."\n\nIn a broken process, the interviewer listens to the story and writes "handled it well" in the Greenhouse or Lever feedback form. In a rigorous process, the interviewer compares the response to a rubric. \n\n## Building a five point behavioral scale\n\nA functional rubric uses a 1 to 5 scale where each number corresponds to specific behaviors, not just a feeling of "poor" or "excellent."\n\nFor a score of 1 (Unsatisfactory), the candidate might have blamed others, ignored the conflict, or allowed the project to fail without communicating the issue. \n\nFor a score of 3 (Proficient), the candidate identifies the conflict early, schedules a meeting to discuss the timeline, and reaches a compromise that keeps the project on track. \n\nFor a score of 5 (Exceptional), the candidate not only resolves the immediate conflict but implements a new process, such as a shared RAID log or a weekly status report, to prevent the same friction from happening again. \n\nBy defining these levels before the first resume is even screened, you force the hiring manager to decide what quality looks like in the context of their specific team. This prevents the goalposts from moving once you meet a candidate who is personally charming but technically average.\n\n## The problem with consensus debriefs\n\nOne of the most common mistakes in mid-sized companies is the group debrief where everyone shares their opinion at once. This leads to "groupthink," where the most senior person in the room or the most vocal interviewer influences everyone else. \n\nTo maintain the integrity of your structured scoring, implement a "blind feedback" rule. No interviewer should be able to see the scores or notes of another interviewer until they have submitted their own. Tools like Ashby or Workable have settings to enforce this. \n\nOnce the feedback is submitted, the recruiter should look for outliers. If Interviewer A gave a 1 on technical skill and Interviewer B gave a 5, the debrief should focus exclusively on that discrepancy. Ask them to point to the specific evidence in their notes. Did the candidate provide a different example to each person? Or is one interviewer applying a higher standard than the other? This level of detail is only possible when you have a rubric to refer back to.\n\n## Training the hiring team to use rubrics\n\nYou cannot simply email a PDF of a rubric to a busy VP of Engineering and expect them to use it correctly. You need to conduct a brief calibration session for every new role. \n\nGather the interview panel for 30 minutes. Review the rubric and ask everyone to describe a hypothetical "3" answer versus a "5" answer. This aligns the team on the bar for excellence. In Europe, where notice periods can be three months or longer, the cost of a bad hire is exceptionally high. Spending 30 minutes on calibration is a cheap insurance policy against a 50,000 Euro mistake.\n\nDuring the interview itself, instruct your team to use the "STARR" method (Situation, Task, Action, Result, Reflection) for their own note-taking. The "Reflection" piece is often missed. A candidate who can explain what they would do differently next time is showing a level of maturity that usually maps to a 4 or 5 on a behavioral rubric. If your interviewers are not writing down what the candidate actually said, they are not grading against the rubric; they are grading against their memory, which is famously unreliable.\n\n## Eliminating culture fit as a metric\n\nIf you see the term "culture fit" in your interview rubrics, delete it immediately. It is a bucket for unconscious bias. Instead, break culture down into specific observable behaviors that reflect your company values. \n\nIf your company values "radical transparency," the rubric should look for evidence of the candidate delivering difficult feedback or admitting to a failure. If you value "frugality," look for examples of resourcefulness under tight budgets. By turning abstract values into scored behaviors, you make culture something you can measure objectively. \n\nThis shift is particularly important for companies growing from 50 to 500 people. At 50 people, the founder usually decides who fits. At 500, that is impossible. You need a system that allows a middle manager in a different country to make the same quality of decision the founder would make, without the founder being in the room.\n\n## Data over intuition\n\nThe final step in fixing the scoring gap is reviewing your data every quarter. Look at your hires who are high performers after six months. Did they actually score higher in the interview process? If your top performers all scored a 3 on "technical agility" but a 5 on "problem solving," your rubrics are telling you something about what actually drives success in your organization. \n\nYou can then adjust your hiring process to weight those specific competencies more heavily. This turns recruiting from a reactive function into a scientific one. It moves the conversation away from "I have a good feeling about this person" and toward "This candidate has a 90 percent probability of meeting their KPIs based on our historical scoring data."\n\n## Implementation checklist\n\nTo start this next week, follow these steps for your most critical open role: \n\n1. Identify the top four competencies required for the job. \n2. Write one behavioral question for each competency. \n3. Define exactly what a 1, 3, and 5 response looks like for each. \n4. Update your ATS to include these specific rubrics in the feedback form. \n5. Require all interviewers to submit scores before the group debrief. \n\nBy closing the scoring gap, you ensure that the structure you worked hard to build actually dictates the hire. You will find that your hiring decisions become faster, your team becomes more diverse, and your retention rates improve because you are finally hiring for the right reasons." }

Sources

  1. 01Structured Interviewing: How to Design and Conduct Effective InterviewsU.S. Office of Personnel Management
  2. 02Interviewer Assessments: A Guide to More Effective InterviewingSociety for Human Resource Management (SHRM)
  3. 03Selection Errors and How to Prevent ThemHarvard Business Review
  4. 04The Science of Hiring: A Review of Structured vs. Unstructured InterviewsJournal of Applied Psychology
ShareLinkedInXEmail

Read next in hiring process

The newsletter

One edition roughly every two weeks: new articles, and what changed in hiring that is worth your time.

Back to all articles