Should confidence scores feed into performance reviews?

No — and the reason isn't squeamishness, it's that doing so destroys the measurement.

Once a measure becomes a target it stops measuring what it used to. Attach consequences to a confidence score and the rational response is to report confidence regardless of what you believe. That takes two or three cycles. After it, the score is uniformly high, the projects slip exactly as before, and you've lost the one channel that was telling you the truth.

The concern is not hypothetical

It was raised, unprompted, by several of the engineering leaders we interviewed — and notably by leaders rather than by the people who'd be measured:

"What if very senior management used this as performance management?"
"Does this become a hammer in the hands of management?"

And from a developer, describing the effect on his own behaviour with no accusation attached:

"No matter how innocuous it is — if this is being used by my boss, it will colour my feedback."

That last quote is the whole argument in one line. He isn't saying he'd lie. He's describing what everyone does when an assessment they provide has consequences for them.

The mechanism, step by step

This has a name. Goodhart's law: when a measure becomes a target, it ceases to be a good measure. The sequence with a confidence score is predictable enough to write down in advance.

Cycle one. A team reports a run of low confidence. In the review, a manager mentions it — possibly not critically, possibly just as an observation. Word travels, because it always does.

Cycle two. The team now knows the number is read as a statement about them rather than about the project. Individuals begin rounding up. Not dishonestly — a 3 becomes a 4 because "we'll probably sort it out," and that's a defensible reading.

Cycle three. Rounding up is now the norm. The team's scores are consistently high. Anyone reporting a genuine 2 is now visibly out of step with colleagues, which makes honesty socially expensive as well as professionally risky.

Steady state. Scores sit between 4 and 5 permanently. The deadline slips arrive with exactly the same lack of warning as before. The dashboard is green, and the organisation is more confident than it was, which is worse than having no dashboard at all.

Nobody at any stage did anything unreasonable. That's what makes it reliable.

Why this is worse than losing an ordinary metric

Two reasons.

The signal was voluntary. Most engineering metrics — deploy frequency, cycle time, incident counts — are byproducts of work happening. They're gameable but not withholdable. A confidence score is an opinion someone chooses to give you, so it has no floor. It doesn't degrade gracefully; it stops.

Trust doesn't rebuild on the same timescale. Undoing this means convincing people that the data won't be used against them, having already used it against them once. In practice that means a long period during which the numbers are meaningless and you can't tell whether they've become meaningful again.

Where the line actually falls

Some uses are clearly fine and some clearly aren't. The middle is where people get into trouble.

Legitimate: deciding where to spend your attention. Prompting a conversation with a team lead. Deciding whether to move the date, cut scope, or add help. Noticing that three teams depending on the same platform all declined at once. Reviewing, after a project ships, whether the signal moved before the problem surfaced.

Not legitimate: anything that treats the number as an assessment of a person or a team. A team lead's calibration discussion. Comparing teams and asking why one is lower. Setting a target for the score. Asking a team to explain their number as though it were a result they produced.

The grey area — and the one to watch: "we should talk about why your team's confidence is low." That sounds supportive, and sometimes is. But it puts the team lead in the position of defending a number they didn't individually produce, and the fastest way to make that conversation go away next quarter is to have a higher number. The safer version asks about the work: what's in the way, what would help. Same conversation, and it doesn't teach anyone to manage the metric.

The two-part test

Before using confidence data in any context, ask:

What to do instead

If you need a measure of a team's effectiveness, use something that isn't an opinion the team chooses to give you: delivery against commitments, cycle time, quality signals. Those have their own problems — every engineering metric does — but they don't collapse the moment they're observed, because they're byproducts rather than statements.

And keep the confidence signal for what it's for: telling you where to look, early enough to act. A leader we spoke to described the posture that keeps it alive:

"You've got to encourage the ugly. Not green is OK."

What we can and can't do about this. Genchi doesn't provide per-person data, individual histories or participation leaderboards — not hidden, not built. That prevents the individual-level misuse. It cannot prevent a leader treating a team's aggregate as a verdict on that team. No vendor can. This is a question about how you lead, and the tool inherits your answer.

A signal built to stay honest

No per-person data. No individual history. Aggregated before anyone sees it.

SEE HOW THE DATA WORKS

Or read our security policy.