Vol. IIIIssue 35Thursday
The Briefing
← Back to all reviews
Productivity ToolsThe Review

Calibration Meetings: How Managers Keep Reviews From Drifting Apart

Two managers, two different rating habits, two very different outcomes for equally strong work. How calibration meetings pull reviews toward a shared standard — and how they go wrong.

Aug 28, 20260.0 / 5
Calibration Meetings: How Managers Keep Reviews From Drifting Apart
Photograph for BusinessWeekly Pro.

In this review

  1. What actually happens in the room
  2. Why it fails when it becomes horse-trading
  3. The role of the facilitator
  4. What changes for the employee, and what shouldn't
Editorial Scoring · Calibration Meetings
CriterionScore
Editorial Score0.0
Value for Money2.0
Implementation Effort2.0
Vendor Trajectory2.0
Overall1.50 / 5.00
Above the fold

Two employees do comparably strong work on comparably sized teams. One gets a review that says "exceeds expectations" and a meaningful raise. The other gets "meets expectations" and a standard increase. Nothing about their actual performance explains the gap — what explains it is that they have different managers, and the two managers rate differently. One is a tough grader who reserves the top rating for something exceptional. The other hands it out generously to keep morale up. Multiply this across a company with dozens of managers and the review system stops measuring performance and starts measuring who you happened to report to. This is the problem calibration meetings exist to solve, and understanding the mechanism explains both why they work and why they so often go badly.

What actually happens in the room

A calibration meeting brings together a group of managers — usually peers within a department, sometimes cross-functional — along with their own manager or an HR partner, to review draft ratings before anything is finalized or communicated to employees. Each manager has already written their initial assessments. The meeting's job is not to write reviews from scratch; it's to look at the distribution across managers and ask whether it holds together. If one manager's team is entirely rated at the top and another's entirely in the middle, calibration is where that gets examined — not necessarily overturned, but examined, with the manager asked to defend the pattern against specific examples.

The mechanism that makes this work is comparison at the point of judgment, not after. A rating written in isolation reflects one manager's internal scale, built from their own career, their own standards, and often their own personality — generous or exacting by temperament, independent of the team they happen to manage. A rating discussed alongside five or six peer managers' ratings, with specific examples on the table, gets pulled toward a shared standard because outlying judgments have to be justified out loud, in front of people who will ask why.

Why it fails when it becomes horse-trading

The most common way calibration goes wrong is that it stops being about accuracy and becomes about protecting headcount. Every manager in the room knows the company has an informal or explicit expectation about the shape of the rating distribution, and if ratings feed into a limited compensation pool, every manager also knows a rating bumped up for someone on their team pulls the pool from somewhere else. The natural response is to advocate hard for your own people and stay quiet about everyone else's, which turns the meeting into negotiation rather than calibration. The manager who argues loudest and longest gets their team rated well; the manager who's conflict-averse or newer to the room gets steamrolled, and their team's ratings suffer for reasons that have nothing to do with the work.

The fix isn't eliminating advocacy — managers should advocate for their people, that's part of the job. It's requiring that advocacy be evidence-based and specific. A calibration meeting with a norm of "bring examples, not adjectives" behaves very differently from one where managers are allowed to argue in generalities. "She's one of my strongest performers" is not evidence. "She redesigned the onboarding flow that cut support tickets by a visible amount and mentored two junior hires through their first quarter" is something the room can actually weigh against a similar claim from another manager.

The role of the facilitator

Calibration meetings run well or badly largely based on who's running them. A weak facilitator lets the loudest managers dominate, lets the meeting run long without reaching resolution, or lets it become a forced-distribution exercise where ratings get shuffled to hit a target curve regardless of whether the curve reflects reality. A strong facilitator does a few specific things: sets the evidence-based ground rule explicitly at the start, keeps a visible tally of where ratings are landing across the group so patterns are obvious rather than felt, and is willing to table a disputed case rather than force a decision in the room when the disagreement reflects a real information gap — sometimes the honest outcome of calibration is "let's get more input before finalizing this one," not a forced resolution in the meeting.

The facilitator also has to protect against a subtler failure: reviewers anchoring on the first few cases discussed and rating everything after relative to those, rather than each case on its own merits. Rotating the order employees are discussed in, or grouping by role rather than by manager, helps counter this drift.

What changes for the employee, and what shouldn't

Done well, calibration changes the number an employee sees without changing the substance of what a manager tells them in the actual review conversation — the manager still delivers feedback grounded in real examples, still has the direct conversation about strengths and growth areas, and the calibrated rating simply reflects that the number attached to that conversation now means the same thing across the company. Done badly, calibration produces a rating that contradicts what the manager has been telling the employee all along, because the number moved in a room the employee wasn't in and the manager doesn't fully explain why.

The practical safeguard is a rule that's easy to state and often skipped: if calibration changes a rating meaningfully from what the manager originally drafted, the manager owes the employee a clear, specific explanation of what changed and why — not "it got calibrated," which employees correctly read as evasion, but the actual reasoning, translated into terms about the work. A calibration process employees can't see the logic of erodes trust in performance reviews faster than almost anything else a company does, even when the underlying process is fairer than what it replaced.

Below the fold · The bottom line
CommentsReader Reactions (0)

Be the first to add to the record.

Letters to the Editor

Leave a comment.

First-time commenters are moderated. Stay on topic. Disagree freely — we publish dissent.

Email is not published.

The Weekly Briefing

Did this review help?

Get one of these on your desk every Monday morning. Free, opinionated — includes clearly marked offers from our partners.

MoreRelated on the Productivity Tools desk