A supplier scorecard is useful when it helps someone decide what to do next. It becomes frustrating when a supplier gets a red rating and nobody can explain the denominator, the source records, or the reason the rating changed.
Build the scorecard around questions your team can answer with evidence: Are commitments slipping? Which parts are exposed? What is missing? Who will follow up? This guide includes an ungated CSV template you can open in a spreadsheet and adapt. Acme and Northfield are fictional, and all example values are illustrative.
Choose the decisions the scorecard should support
For a weekly operational review, the decision might be which suppliers need a call and which commitments need a recovery plan. For a quarterly business review, it might be which recurring causes need joint improvement work. These reviews can share evidence while using different time windows and levels of detail.
Avoid a single blended number until your team can explain its components. A high delivery rating should not hide an unresolved quality issue, and a small spend share should not hide a critical sole-source part. Keep the measures separate so the reason for attention remains visible.
Agree who owns each measure. Procurement may own commitments and supplier follow-up, receiving may own receipt completeness, and quality may own accepted defect records. Name the owner of missing data too; a scorecard cannot become reliable if every gap is someone else’s responsibility.
Use one row per supplier, measure and review period
The downloadable template is deliberately in a long format: one row for each measure. It includes supplier and site IDs, period dates, measure definition, numerator, denominator, value, unit, evidence reference, data coverage, review status, reviewer, action owner, action due date, closure criteria, target, target effective date and next review date.
This structure keeps a delivery percentage and a count of overdue lines from being mistaken for comparable scores. Filter it by supplier for a review meeting, or by owner for follow-up. The template contains one fictional filled-in delivery example and blank rows for all five measures. It is a starting worksheet, not an automated rating engine.
| Measure | Definition | Evidence needed |
|---|---|---|
| On-time delivery | Full receipt by agreed commitment / eligible completed lines | PO lines, commitment history, receipts |
| Open overdue lines | Unfulfilled eligible lines past commitment at cutoff | Open quantity and cutoff timestamp |
| Lead-time change | Recent actual lead time compared with stated baseline | Order and completion dates, comparable part mix |
| Quality exceptions | Accepted quality events under a stated counting policy | Inspection or defect records, affected quantities |
| Data coverage | Eligible records with required fields / expected eligible records | Missing-field report and scope definition |
Explain a rating with the underlying calculation
Suppose Northfield’s July review includes twenty completed Acme lines, sixteen received in full by the original promised date. The on-time result is 16 ÷ 20 = 80%. Keep the twenty source-line references available, including the four late ones, and record whether the calculation used the original commitment or an accepted revision.
A team might choose to review any supplier below its agreed target. That target is a business policy, not a universal industry threshold. Record the target and its effective date rather than presenting an amber status as an objective fact. A reviewer should be able to understand the result without reverse-engineering a spreadsheet formula.
Show missing evidence instead of inventing a score
If five of twenty-five expected completed lines have no usable promised date, show twenty scored lines and five excluded lines. Coverage is 20 ÷ 25 = 80%. The delivery result describes the twenty eligible lines, not necessarily all twenty-five. Missing dates can bias the result if the excluded orders are the troublesome ones.
Use labels such as insufficient evidence, awaiting receipt confirmation or reviewed. Do not turn an unknown outcome into a successful delivery. Keep the list of exclusions and their reasons, and assign someone to investigate whether the source export or the underlying records need correction.
Avoid treating a small sample as stable performance. Two of two on-time lines gives 100%, but it does not offer the same basis for judgment as two hundred of two hundred. Include the count and the review period wherever you display a percentage.
Keep performance, exposure and predictions distinct
Past delivery performance describes what happened. Exposure describes what a failure would affect today: a critical part, limited cover, a long replacement lead time or a concentrated production site. A prediction describes an event that may happen within a future window. Putting all three under a column called risk makes the scorecard harder to use.
A supplier can perform well historically while your business is highly exposed to it. A historically late supplier can be less urgent this week if stock cover is healthy. Display enough context for a buyer to make that distinction and explain the decision.
If you include a forecast, show the event definition, horizon, evidence cutoff and evaluation history alongside the output. Do not copy a confidence value into a probability column. Check the definition of each forecast field before interpreting it.
Make the review produce an owner and an action
Begin the meeting with changes since the last review, new exceptions and missing evidence. For each issue, agree the next action, owner, due date and what would close it. “Monitor Acme” is harder to follow up than “confirm the recovery date for the two casting lines by Thursday.”
Keep supplier responses and internal corrections attached to the issue. If receiving confirms that a receipt was posted late, record that correction rather than silently replacing last month’s score. People need to distinguish a change in supplier performance from a change in the evidence.
At the next review, check the action outcome as well as the metric. A recurring late-delivery result may need a different response if the original improvement action was never completed.
Move from a worksheet to connected, reviewable evidence
The template helps establish the definitions before you automate anything. In Semogram, connect the relevant sources, reconcile supplier identities, and build queries over commitments, receipts and evidence. The resulting answer should let the team inspect the source records and the execution that produced it.
Use assertions and review records to preserve claims, supporting or contradicting evidence, and corrections. Inspect whether the result includes the records needed to explain each score. Investigate missing evidence before bringing the score to a review meeting.
You can then ask a bounded question such as “Which suppliers changed materially this month, and which records explain the change?” Review the answer and decide the next action. Forecasting can add a future-looking measure, but it needs its own evaluation rather than borrowing credibility from a well-defined historical scorecard.
Apply this to your own operations
Start with one question and the records behind it. We can help you scope a workflow your team can inspect, correct and evaluate.