Atmospheric dark artwork of a bank of analogue gauges frozen mid-reading, a single dial picked out in editorial blue

📍 IN BRIEF A vendor scorecard is a contract artefact. It was an accurate picture of the service on the day the deal was signed, and it has been drifting away from reality ever since, because the platform changes every quarter and the metrics change at renewal. The real governance work is not chasing the scores. It is re-choosing the metrics on a dated cadence, with an owner and a retirement path for every line.

Most vendor scorecards I see have been reviewed dozens of times and changed never. The quarterly meeting runs, the numbers get read, actions get logged, and the list of metrics itself sits untouched from one contract year to the next. The service being measured has changed shape several times in that period. The instrument measuring it has not.

Before I read a single number on a supplier scorecard, I have learned to ask a different question, which is “when was each of these metrics last added, challenged or removed”. In most of the reviews I have sat through, the honest answer is the day the contract was signed, for every line on the page.

The vendor is honest, the numbers are accurate, and nobody is gaming anything. And the relationship is still deteriorating, because the scorecard is faithfully reporting on a service as it was defined three years ago, while everyone in the room knows the real work has moved somewhere the metrics cannot see.

The scorecard is a photograph of the negotiation

A vendor KPI set records what worried you on signing day, not what the service does today.

Think about when vendor metrics actually get written. They are drafted during the negotiation, under negotiation pressures. Whatever the two sides argued about hardest gets a metric, whatever felt safe gets ignored, and the whole set is frozen into a schedule the moment the signatures land. That set then governs a live service for three years or more, on a ServiceNow estate that will take new workflows, new user groups and new demand every quarter of those years. Good delivery discipline says measures should be selected with the people closest to the process and validated with the governance and finance functions who will steer by them. That validation is usually done once, at the start, and almost never again, which means the scorecard is not a picture of the service. It is a photograph of the negotiation, ageing in its frame.

The British rail industry spent years learning this in public. Train punctuality was long reported as arrival at the final destination within a few minutes of schedule, a measure that suited the network it was designed for. As journeys and usage changed, the gap between that number and what passengers actually experienced at intermediate stations grew wide enough to become a credibility problem, and the industry eventually moved to recording arrival at every stop. Notice what the fix was, not a tougher target on the old measure but a retirement, and a replacement that matched how the service was now being used.

⚠️ COMMON PITFALL. Tightening the target on a stale metric feels like rigour, but it optimises the vendor ever harder against a definition of the service that no longer matches how the service is used.

Green numbers are self-sealing

A stale metric survives precisely because it passes, since review attention only flows to red.

There is a structural reason nobody notices the drift, which is that in any performance review, scrutiny follows the exceptions. A red number triggers questions, evidence, an improvement plan. A green number gets a nod, and the meeting moves on. So a metric that has quietly decayed into irrelevance is protected by its own colour, because the one thing a green line never invites is the question of whether it still deserves its place. The economists' warning applies, usually filed under Goodhart's law, that a measure which becomes a target stops measuring what it was meant to. What gets less attention is the slower version of the same disease, where the measure stops describing the thing at all and keeps passing anyway.

Stale metrics also acquire defenders, because once a number has lived in the contract schedule, the bonus plan and the board pack for a few years, proposing to retire it sounds like proposing to lower standards, and nobody wants to be the person who argued for measuring less. England's four-hour emergency care standard showed how strong that lock-in can be. For the best part of two decades it organised hospital behaviour around a clock, produced well-documented distortions at the edge of the window, and was defended long after clinicians had begun arguing openly that it no longer described good care. Whatever view you take of the eventual reforms, the governance lesson stands. The measure outlived the conditions it was designed for by years, in plain sight, because standing metrics have constituencies and no scheduled exit.

That is how a scorecard reduces in value. Line by line, each one once earned its place, each one now passes automatically, and collectively they give the review a feeling of control that the numbers steering the room no longer justify.

Give every metric a passport

Treat every vendor KPI as perishable, with an owner, a linked outcome, a review date and a retirement path.

The move that changes this is small and unglamorous. Stop treating the metric set as part of the contract's furniture and start governing it as its own asset. Every vendor metric has a half-life, and it is shorter than your contract. So every line on the scorecard carries four things. A named owner on your side of the relationship, because a metric nobody owns is a metric nobody will ever retire. The outcome it is evidence for, stated in one sentence, so the line can be challenged against something. A review date, at most a year out, when the owner must argue for its place again. And a retirement condition, agreed in advance, so removing it is the execution of a plan rather than an admission of defeat.

Then give the quarterly review one standing agenda item alongside the scores. Which line on this scorecard no longer changes any decision we make, and what should replace it? The delivery methodologies most platform programmes already run on point exactly this way, they treat staleness itself as a signal worth tracking for governing documents, and their health-check discipline re-tests the assumptions underneath a benefit rather than just reading the results. Vendor metrics deserve the same suspicion. A service level that breaks at least tells you something real. A service level that can no longer break tells you nothing, and costs you attention every quarter it survives.

The cost of skipping this is not hypothetical. Wells Fargo ran its retail bank for years on a cross-selling number that had long stopped evidencing the outcome it was meant to stand for, customer relationships deepening. The target stayed, the behaviour organised itself around the target, and the affair ended with regulators involved and the bank abolishing product sales goals outright. It is remembered as a fraud scandal, but to a governance eye it is also the most expensive un-retired KPI in recent corporate memory, and the sobering part is that removing the metric was always available, years earlier, for nothing.

Chart comparing two lines across a three-year contract term: a metric set re-validated quarterly saws back up at every review, while a metric set fixed until renewal decays steadily from signature day, with the widening gap labelled as the drift the review never sees

A scorecard reviewed only at renewal spends most of the contract measuring the last negotiation. Illustrative, drawn from platform programme reviews.

Where this doesn't apply

Some services do not drift enough to need this. Where the activity genuinely is the outcome, a leased line staying up, a payroll run completing, conformance metrics can stand for years and a re-validation cadence is ceremony. Short engagements that end inside a year will usually finish before drift costs you anything. And the discipline has an opposite failure worth naming. A scorecard rewritten every month cannot show a trend, and baselines only mean something if the measure underneath them holds still. Retire metrics in ones and twos with their history archived, on the dated cadence, not in wholesale resets that erase your own memory.

The bottom line

Put the scorecard itself on the governance agenda, this quarter, as a decision rather than a discussion. Name an owner for the metric set and date its next review. Ask of every line whether it changed any decision in the last two quarters, and retire the ones that did not, with their replacement pointed at an outcome someone can state in a sentence. Vendor accountability still has to live inside your own house, whatever the supplier runs, and that starts with owning the instrument's contents, not just reading its results.

Your vendor scorecard was accurate the day the contract was signed. Unless someone's job is to re-check it, everything since is drift.

P.S. Pull up your oldest vendor scorecard and find the newest metric on it. If nothing has been added or retired since signature, you have just found the work.