AWRA OpsHub Search

Scoring Suppliers on Evidence: What a Vendor Scorecard Measures

Everyone knows which supplier is difficult. Almost nobody can prove it, which is why the difficult supplier keeps winning orders. What a scorecard actually measures, the weighting behind the number, and the one definition inside it you must check before quoting it at anybody.

AI & Insights Washingtone Aura 11 min read

Ask a procurement officer which of their suppliers is unreliable and you will get an immediate, confident answer. Ask them to demonstrate it and the conversation slows down, because the evidence is spread across forty purchase orders, a WhatsApp thread and the memory of a delivery that arrived nine days late in April.

That gap is expensive. Without evidence, the argument for dropping a supplier is one person's impression against a long relationship, and impressions lose that argument. A scorecard exists to convert what everybody already suspects into something you can put in front of the supplier.

Three things worth measuring

A scorecard that measures everything measures nothing, because a score built from twelve factors cannot be explained to the supplier it judges. Ours uses three, weighted, and the weighting is deliberate.

Factor Weight What it actually counts
On-time delivery 50% The share of this supplier's orders that were delivered within the expected window
Responsiveness 30% The share of orders the supplier acknowledged at all
Consistency 20% How much their delivery lead times vary around their own average

Responsiveness surprises people, and it should not. A supplier who acknowledges an order tells you the order was received, understood and accepted — which is the difference between a delivery you are waiting for and a delivery that was never going to come. Silence is the most under-recorded risk signal in procurement, and it costs nothing to measure.

Consistency is the quiet one. A supplier who always takes eleven days is more valuable than one who averages seven but ranges from two to twenty, because you can plan around the first and cannot plan around the second. That is why variance is scored separately from the average and rated in plain words — excellent, stable, or volatile.

A supplier who always takes eleven days is worth more than one who averages seven and ranges from two to twenty. You can plan around the first. The second makes every one of your own promises a guess.

The definition you must check

Here is the part most vendors would leave out of an article like this. "On time" has to be measured against something, and what it is measured against decides whether the whole score means anything.

In our current implementation, on-time means delivered within a standard window after the order was raised — a uniform expectation applied across suppliers, rather than a comparison against a date each supplier individually promised on each order. Read that twice before you use a score in a negotiation.

What that means in practice — and what to do about it

A supplier whose goods genuinely take three weeks by the nature of the item

The consequence

Scores poorly on on-time even when they hit every date they promised you.

How to work with it

Compare like with like. Rank suppliers of the same category against each other, not a fabricator against a stationery shop.

A supplier who quotes a long lead time and then beats it

The consequence

Gets no credit for the accuracy of their own promise.

How to work with it

Read the average lead time and the consistency rating alongside the score, rather than the score alone.

A supplier whose delivery date on the order was never filled in

The consequence

Cannot be assessed on delivery at all — they simply have less evidence.

How to work with it

Fix the record discipline first. A scorecard is a mirror of your data entry before it is a mirror of your suppliers.

A supplier with two orders in total

The consequence

Produces a score that looks as authoritative as one built from two hundred.

How to work with it

Check the order count on every scorecard before you quote the number. Small samples are the main way scorecards embarrass people.

We would rather you knew that than discovered it in a supplier meeting. Used as a relative ranking within a comparable category, the score is genuinely useful. Used as an absolute verdict on a supplier's professionalism, it will occasionally be unfair — and the supplier will be able to say so.

The concentration question nobody asks

Supplier risk is not only about how well each one performs. It is also about how much of you depends on any one of them, and that is a different calculation entirely — one that looks at the distribution of your spend rather than the behaviour of any individual vendor.

A large share of total spend concentrated on one supplier is flagged as concentration risk, alongside any suppliers whose reliability is deteriorating. Both are worth reading together at least quarterly: excellent performance from a supplier who holds most of your spend is a strength today and an exposure the moment they raise prices, lose a key person or fail.

Sole-sourcing is a decision, not an accident

Most concentration is not chosen. It accumulates because one supplier is easy to work with, so they get the next order, and the next. That is a perfectly good reason to use them — and a perfectly bad reason to have no alternative priced and prequalified. The point of the flag is to make the accumulation visible while you still have options.

Scoring a supplier you have never used

Performance scoring needs history, which leaves the hardest case unanswered: a new supplier with no orders behind them. That is what prequalification is for, and it has its own, deliberately separate, form of assessment.

Where an applicant submits a prequalification application, an advisory evaluation can score the submitted information and flag risks for the reviewer — completeness, consistency, the shape of what was declared. Three things about it should be stated plainly:

  • It reads the typed application data and document completeness, not the contents of uploaded certificates. Whether a tax compliance certificate is genuine and current is not something it can see.
  • It is advisory only. A human approves or rejects, always.
  • It fails soft. Where no AI provider is configured or the request fails, the reviewer simply sees no evaluation rather than a blocked application.

That is a deliberately modest claim. Any vendor telling you their system vets suppliers is describing document verification against issuing authorities, which is a different product and, in most markets, a manual process with a phone call in it.

Making the scorecard change something

  1. Review quarterly, by category

    Fifteen minutes with the category buyer. Same-category suppliers side by side, order counts visible, outliers with tiny samples set aside.

  2. Show the supplier their own numbers

    This is where the value is. A supplier shown "eleven of your last twenty deliveries were late and you acknowledged four orders out of twenty" responds very differently to one told they are unreliable. The first is a conversation, the second is an insult.

  3. Attach a consequence to the pattern, not the incident

    One late delivery is life. A quarter of lateness with no acknowledgement is a pattern, and a pattern justifies moving volume, requesting a discount, or requiring a second source.

  4. Re-score after the conversation

    A supplier who improves after being shown the data is the outcome you actually wanted. Most do improve, which is the argument for showing them.

What we do and do not do

Supplier scoring — the straight answer

What AWRA OpsHub does today

  • A weighted reliability score per vendor — on-time 50%, acknowledgement 30%, consistency 20%, clamped to 0–100.
  • Average lead time and a consistency rating derived from the variance of that supplier's own deliveries.
  • Order counts on every scorecard, so a two-order sample is visible as one.
  • Spend concentration risk across your vendor base, plus flags on suppliers whose reliability is declining.
  • Advisory AI evaluation of prequalification applications, scoring the typed submission and flagging risks for a human reviewer.
  • A plain-language narrative over vendor performance where an AI provider is configured.

What it does not do

  • On-time is not measured against a per-order promised date. It is a standard window from when the order was raised — which is fair between comparable suppliers and unfair between different categories.
  • No quality or defect scoring from receipts. Goods-received discrepancies are recorded in receiving; they do not feed a quality dimension in the score.
  • No price competitiveness in the score. Quote comparison is a separate exercise with its own weights for price, delivery and distance.
  • No document verification. Nothing checks a certificate against the issuing authority, and the prequalification evaluation does not read uploaded files at all.

Use the score to rank suppliers of the same kind against each other. Do not use it to compare a machinery importer with a stationery vendor, and always look at the order count before quoting a number to anyone.

Our take

The scorecard's real value is not the number — it is that the conversation with a supplier stops being about impressions. Review by category quarterly, check the order count, read consistency alongside the score, and show suppliers their own figures. That last step changes behaviour more reliably than any procurement policy we have seen.

See supplier collaboration and vendor performance

Reliability scoring from your own order history, lead-time consistency, concentration risk and advisory evaluation of new supplier applications.

Explore vendor performance

Frequently asked questions

How many orders does a supplier need before the score means anything?

Enough that one bad delivery cannot dominate — as a rule of thumb, ten or more, and read the consistency rating rather than the headline below that. Two-order scorecards look exactly as authoritative as two-hundred-order ones, which is the main way a scorecard embarrasses whoever quotes it. The order count is shown on every card precisely so you check it first.

Our best supplier scores badly. Why?

Almost always because their category naturally takes longer than the standard window used to judge on-time delivery, so an importer of specialised equipment is being measured on the same expectation as a local consumables vendor. Compare within a category rather than across your whole vendor base, and read their average lead time and consistency: a supplier who always takes twenty-one days and never varies is performing well, whatever the on-time percentage says.

Does the score include price?

No, deliberately. Price is decided during quotation comparison, where you can weight price against delivery and distance for that specific requirement. Reliability is a property of the relationship over time, and mixing the two produces a number that says nothing useful about either — a cheap supplier who never delivers on time and an expensive one who always does would land in the same place.

Can the AI evaluation replace our prequalification committee?

No, and it is not built to. It scores the typed application and document completeness to give a reviewer a starting point and some risk flags; it cannot read the contents of uploaded certificates, cannot verify anything against an issuing authority, and fails soft to no evaluation when a provider is unavailable. A human approves or rejects every application. Treat it as a first pass that saves reading time, not as a decision.

What should we do about concentration risk?

Not necessarily change supplier — often the right response is to keep the relationship and remove the dependency, by prequalifying and price-testing a genuine alternative before you need one. The cost of concentration is not paid while things go well; it is paid on the day that supplier raises prices, loses their key person or fails, and you discover that switching takes six weeks you do not have. Review the distribution quarterly and treat any single supplier holding a large share of spend as a standing item.

Help Center

Need a quick answer while you read?

Run inventory, procurement, assets, sales, and field work with approved AWRA guidance for setup, migration, integrations, security, pricing, and support.

Search all approved AWRA public help articles.

Open Help Center