SpacesE · Appraisals, goals and succession

Calibration: getting two managers to mean the same thing.

A calibration session compares ratings before they are published. What it changes, what it must record, and why the manager's original score stays on the file.

The XUnframed Team
5 min read

Calibration is a meeting where managers compare draft ratings across teams before anything is published, so that a 4 in one department means roughly what a 4 means in another. The output is concrete: a changed score against a named review, a written rationale, and the people who agreed it.

It happens between the manager review phase and publication. That position in the cycle is the whole design. After publication the same conversation is an appeal, and an appeal carries a heavier burden for everyone in the room.

The problem calibration actually solves

Two managers with identical teams produce different rating distributions, and both believe they are being fair. One has an internal bar set by the strongest person they ever managed. The other rates against the job description. Neither is behaving badly, and no amount of guidance in the review form fixes it, because the difference lives in how each of them reads the scale.

The effect compounds when ratings feed pay. Acas notes that the process for deciding performance-related pay should be fair and objective, that employers must avoid treating people less favourably because of a protected characteristic, and that a system people see as unfair leaves them less motivated and more likely to leave.

What a session has to record

Our calibration sessions belong to a cycle and carry a named list of participants, a status of draft, in progress or finalized, and a list of adjustments. Each adjustment names the review it changes, the calibrated rating, an optional potential rating, and a rationale that is required rather than optional.

Requiring the rationale is the part that changes behaviour. A session where six ratings moved and none of the moves has a sentence attached produced a different distribution and no evidence, and it is indistinguishable afterwards from a session that never happened.

The calibrated rating sits next to the manager's own rating on the review rather than replacing it. A reviewer's original score, the agreed score and the reason for the gap all stay on the file, which is what lets someone answer a question about that review two years later.

Blind mode, and when it helps

A session can run in blind mode, with identifying detail hidden while the evidence is compared. It is useful for a first pass across a large population and it has a limit worth naming: the evidence itself often identifies the person to anyone who works with them. Treat it as a way to slow down recognition rather than to remove it.

Low scores need evidence before they need a meeting

Calibration works on what the reviews contain, so the quality of the inputs decides the quality of the session.

We set a negative threshold as a percentage of the scale, and any score below it requires structured evidence before it can be submitted: the situation, the behaviour, the impact and a specific example. That turns a 2 out of 5 from an opinion into something a calibration session can examine.

A score below the threshold with evidence attached also opens a flag in an HR queue. HR can uphold it or reject it, and a rejected score stops counting towards the overall rating rather than being deleted. A daily reminder chases flags that have been sitting past their service level, because the queue that nobody is chased about is the queue that holds up publication.

Running the meeting

Three practical rules do most of the work.

Bring the evidence, not the distribution. Open with the reviews that sit at the ends of the scale and the ones where the self assessment and the manager assessment diverge most. A session that starts with a curve spends its time defending the curve.

Give every manager the same slot. Otherwise the loudest advocate calibrates the whole cohort upwards.

Record the decision in the room. A rationale written the following week is a reconstruction.

The line the session must not cross

Calibration changes a rating. It does not, on its own, decide pay or an exit. Those decisions stay with a named person who can explain them, and the rating is one input.

The ICO sets out that a solely automated decision producing legal or similarly significant effects is only permitted in narrow circumstances, and that people must be able to obtain human intervention, express their point of view, get an explanation and challenge the decision. A calibrated score that is wired straight to a pay outcome with nobody in between is precisely the arrangement that guidance addresses. Keep the human step, and keep the record that shows who took it.

Acas adds the documentation duty from the other end: keep a written record of what was discussed in a review and share it with the employee afterwards.

What good looks like afterwards

Count three things when the session closes. How many ratings moved. How many of those moves have a rationale long enough to be useful to someone who was not in the room. And how many reviews were discussed without changing, which is the number that tells you the session was examining evidence rather than hunting for adjustments.

How to run an appraisal cycle end to end sets out where this phase sits and what has to close before it opens. Performance management without a separate tool covers what happens to the calibrated rating once pay decisions start reading it.

Questions

Answered here.

What is performance review calibration?
It is a meeting where managers compare draft ratings across teams before results are published, so that the same score means roughly the same thing in engineering as it does in finance. The output is a changed rating against a named review, with a written rationale and the person who agreed it.
When should calibration happen in the cycle?
After manager reviews are submitted and before results are published. That ordering is what makes it calibration rather than an appeal. Once an employee has seen a rating, changing it is a different conversation with a different burden of explanation.
Does calibration overwrite the manager's rating?
It should not. In Unframed HR the calibrated rating is stored alongside the manager's own, so the file shows the original score, the agreed score and the rationale for the difference. An overwrite destroys the evidence that the session did anything.
Is calibration a forced distribution?
No. A forced distribution fixes how many people may receive each rating before anyone is discussed. Calibration compares the evidence behind ratings and changes individual scores where the evidence does not support them, which can leave the distribution exactly as it was.

Sources