Guides

Calibration without politics

Author

MeritFlo Editorial

Date Published

Calibration without politics — MeritFlo guide cover

Calibration has a deserved bad reputation in many companies: a room where scores go in, horse-trading happens, and different scores come out — with the advantage going to whichever manager argues loudest. But the alternative to bad calibration is not no calibration; it is uncorrected rating cultures, where working for a generous scorer is worth more than performing well. The fix is procedural. Calibration goes wrong when it debates people; it works when it debates distributions.

Start from the shape, not the names

Open the session with each team's score distribution side by side — no names, just shapes. One team averaging 4.4, another 3.1, on comparable work, is a rating-culture gap, and it is visible in thirty seconds. Agree on what the distributions say before discussing any individual. This single sequencing change removes most of the politics: managers defend their scoring standard, not their favourite people.

The evidence rule

One rule transforms the individual discussion: every score defended must be defended from the record. "She is clearly a top performer" is not admissible; "her scores against these criteria, across these reviews, with this peer feedback" is. The rule advantages managers who documented all year over managers who advocate well in meetings — which is exactly the incentive you want to create, because next cycle they all document better.

Calibration smell

What it usually means

The procedural fix

Loudest manager wins

Debate is about people, not standards

Open with distributions, no names

“Clearly a top performer”

Advocacy without evidence

Evidence rule: defend from the record only

Dozens of individual re-scores

Criteria were too vague at design time

Fix criteria upstream, adjust standards here

Employees surprised by changes

Calibration happening in secret

Announce provisional scores at kickoff

Adjust standards, not verdicts

When a team's scores shift in calibration, the adjustment should be a standards correction applied consistently — "this team's 5s are everyone else's 4s" — not a retrial of individuals. Individual re-scoring belongs in exceptional cases where the evidence and score plainly diverge. If the session is re-litigating dozens of individual scores, the criteria were too vague at design time; fix that upstream rather than compensating in the room.

Document what moved and why

Every adjustment needs a recorded reason. This is partly defensive — calibrated scores feed compensation, and compensation decisions get challenged — but mostly cultural: a manager who must write down why a score moved argues differently from one who only has to win a conversation. In MeritFlo, calibration adjustments are recorded alongside the original scores, so the full chain from raw review to final number stays auditable.

Tell employees calibration exists

The most common calibration failure is secrecy. An employee who learns after the fact that "a meeting changed my score" has every reason to distrust the system. The message is simple and honest: scores are provisional until they are checked for consistency across teams, because we do not want your rating to depend on who your manager is. Framed that way, calibration is not the politics — it is the protection against politics.

Run this way, calibration sessions get shorter every cycle. The first one surfaces big rating-culture gaps; by the third, managers have internalised the shared standard and the session mostly confirms it. That trajectory — loud then quiet — is how you know it is working.