Last week, I posted the following on LinkedIn:
Here’s an idea that is sure to get some (many?) folks up in arms:
Promoting on savings delivered = promoting on evidence that is getting automated away from us.
I know, I know, I can see the eye-rolls already, but let me explain what I mean.
The metrics that built every Procurement career over the last thirty years (savings, efficiencies, throughput) describe exactly the work that most AI tools are being built to deliver.
Which means when we promote our “top performers” on the basis of those metrics, we are, in a very real way, promoting someone on an evidence basis that the machines are doing a great job of taking over (at least on a relative cost basis).
So, in that context, performance in the current role (via the metrics above) no longer predicts capability at the next level - and certainly won’t in a post-AI world. AI and the machines are severing that link.
So what does?
Judgement under ambiguity and an orientation towards outcomes, as credited by the actual stakeholders we serve.
Whether we’re building real judgement in others.
Whether our stakeholders want us in key discussions - and much earlier in the process.
Of course, none of the above show up in our usual savings dashboards, which is a problem.
It’s time for us to think more carefully and proactively about our promotion criteria.
What are you promoting on?
In response, a reader asked:
How do you measure “judgement under ambiguity” without turning it into another dashboard people learn to game?
That’s a really great question because it hits on a number of veins - whether such metrics are relevant, whether we can even measure such traits effectively, and whether we can embed them while keeping the entire process honest.
It’s a topic worthy of deeper discussion and one I’ll dive into in this post.
We Keep Choosing Savings
The topic of metrics is one that Procurement and the broader business world has been tackling for years.
Ideas like the Balanced Scorecard have, for decades, pushed forward the idea that we need a diverse set of metrics to truly gauge the full spectrum and complexity of individual performance - and, therefore, that the measurement systems we utilize must encompass both hard and soft metrics.
But despite our best intentions, when push comes to shove, we don’t always practice what we preach. In actuality, we tend, more often than not, to revert back to the same old measures we’ve always relied on: savings delivered, cost avoidance achieved, cycle time reductions, compliance rates, spend managed, etc.
Why is that?
Not Fuzzy - Inconvenient
The usual explanation is that soft metrics are hard to measure. They’re fuzzy, hard to pin down and certainly harder to compare across people at similar levels than say, how much money did you save us last year?
This argument is only partly true - there is some truth to these concerns but the fact is also that a good manager can, given the chance, usually assess capabilities such as judgement fairly consistently and accurately when given the chance.
The real reasons, then, have to do with culture, design and reporting.
Many organizations speak early and often about the value of metric diversity, but when you look beneath the hood, you find that the culture doesn’t truly prize it. It’s not baked into the language of management, into the initiatives they fund, or the reward systems they put in place. At the end of the day, hard numbers matter because incentives (i.e. compensation and promotion) are aligned to them.
Second, the metrics themselves aren’t designed for success - which leaves them open to gaming by individuals who work to the letter of the law rather than the spirit of it. Alternatively, we also see situations where the manager assessing the employee’s performance on these measures simply doesn’t have the ability or experience to be able to make an accurate evaluation.
Last but not least, ‘soft’ metrics tend not to travel well up the chain. When managers have to report upwards on a team member’s performance, they find it harder to sum up such performance and compare them across individuals in a way that is reasonably objective and ‘backable’ - whereas hard metrics don’t have this problem.
The Harder Work
So what are we to do?
The point, in my mind, isn’t to abandon the idea of soft metrics. We absolutely need balanced metrics, especially in a post-AI world, where the machines are taking on much of the “easy to measure” work and where the ‘softer’ metrics matter even more.
We need to actually embrace them even more, and do the ‘harder’ work of doing it right - which means defining them better, teaching managers to evaluate better, and establishing real rewards and consequences.
Describe Reasoning, Not Results
First off, we need fewer metrics that are better defined; that is, each metric has clarity behind what it means and what we expect from someone exhibiting those ‘soft’ behaviors.
The key is to be careful not to turn this into a checklist. The more tightly we try and define it, the more specific we try and make the constituent parameters, the easier it becomes to game.
What matters, instead, is to define what ‘good’ looks like in terms of the metric itself - for example, how does the individual weigh different signals, what did they base their decisions on when there was little data available, what information did they consider more versus less relevant when making a decision.
In this sense, competency models, by level, do a great job. I used these all the time in both my consulting days as well as with The Smart Cube as I assessed my teams. Such models show the nature of the skill at each level as well as the expected trajectory. They allow both the manager and the team member to understand what “good” actually is and work towards it.
Assessing Judgement Requires Judgement
Irrespective of how well we define our metrics or how good our competency models are, the fact is that assessing judgement itself requires judgement.
What does that mean? Well, that means that we, as managers, have to be able to measure the decision-making process of our people rather than (just) the outcomes they’ve produced. I appreciate that this might sound somewhat counter-intuitive but it actually isn’t. Good decisions can sometimes go wrong while bad ones can sometimes get lucky. The point is to keep the focus more towards the reasoning and thought process behind each decision.
Which, of course, raises an obvious problem: an organization can only measure judgement to the extent that its managers actually have that judgement - and the ability to measure it - themselves. That’s not always the case.
Years of incenting hard metrics over soft ones means many managers either never developed this ability or, at the very least, their skills are flabby. This is the apprenticeship problem I’ve discussed before but one level up from that of the junior: many of our managers can’t make this evaluation because they themselves came up on the “savings as the only metric” regime. They don’t know how.
Solving for this needs work - and it isn’t going to be a simple fix.
It’s a multi-year build. It requires managers with the ability to make these judgements and the stomach to navigate this more ambiguous process rather than the formulaic one. It also requires managers calibrating their reasoning and how they evaluate the same behavior or trait; the divergence of viewpoints at such calibration sessions exposes how and where standards differ, allowing for shared understanding to be built.
Make It Count
Lastly, we won’t succeed unless we reward the behaviors we want and penalize the ones we don’t. At the end of the day, people will only do what they’re incented to do. Getting to this has a hard part and a practical part.
The hard part is that the process to make it count is not a simple one. The solution cannot be to establish a formula per se to determine who gets a reward and who doesn’t; we’re not trying to get to a “score”, because that can be gamed. We need, instead, the right manager to conduct an assessment that is thoughtful but still discretionary, one where he or she exercises their own judgement in the process. This is, as I’ve said before, hard work but the benefit is clear: without a score to be achieved, the team member must then understand the required behaviors to model and, then, actually be good.
In terms of practical levers, rewards and consequences matter. Compensation (in the form of a bonus) is certainly helpful in this regard, but promotion and advancement is more meaningful and long term. Advancement within the organization is the clearest signal of what the culture values. Promote the right behaviors even a couple of times, and suddenly everyone will get the message as to what’s important.
So, What’s the Answer?
Look, at the end of the day, I’m not sure you can build an ungameable metric, but you can build metrics that resist it.
Metrics that are finite but diverse. Metrics that are meaningful, encompassing the hard and the soft. Metrics that aren’t about “scores” but about competencies, decision patterns, and good reasoning, judged by strong managers consistently and regularly. Metrics that are rewarded appropriately.
As AI compresses our “visible” outputs and takes on much of our work that is measurable, the stuff that will remain relevant will be the work of the human, the ‘soft’ stuff. We need to learn to manage and reward these more consistently and effectively.
So my answer to the LinkedIn commenter on my post: you cannot measure judgement under ambiguity, you have to assess it, repeatedly, by someone qualified to do so, against a defined set of competencies and a trajectory (versus a score), with real implications (rewards or consequences) attached.



