How to test a supplier-risk prediction before trusting it

Evaluate supplier-risk predictions against what happened, with a worked back-test, false-alarm analysis, evidence cutoffs, calibration and a practical pilot checklist.

By Semogram · 8 minute read

For operations leaders evaluating AI forecasts or a supplier-risk pilot

A forecast can sound persuasive and still be unhelpful. “Acme is at elevated risk” leaves too much unspecified: risk of what, by when, based on which records, and useful for which decision? Before relying on a prediction, make it precise enough to compare with an observable outcome.

This guide shows how to evaluate a bounded supplier-delivery question. All counts and probabilities are fictional examples for explaining the method. They are not Semogram performance claims. Forecast quality must be measured on your own data and under the conditions in which your team will use it.

Define an event your team can actually observe

Choose one event, subject and horizon. For example: “Will supplier Acme have at least one eligible order line miss its original promised date during the next 30 days?” Define eligible lines, receipt completeness, accepted cancellations, tolerance and the site or legal entity being assessed.

For this example, freeze the eligible lines at the prediction cutoff: include uncanceled lines with an outstanding quantity and an original promised date after the cutoff and within the next 30 days. Exclude lines already overdue, and assess newly placed orders in a later prediction. A line misses its date if the required quantity remains unreceived at the end of that promised day. Record later cancellations separately under a written outcome policy; do not silently count them as successful deliveries.

If a supplier has no eligible lines due, mark the delivery forecast not applicable and exclude it from this evaluation. Keep that count visible. Missing promised dates or unreliable open quantities are evidence gaps, not proof that no deliveries are due. Investigate them before assigning eligibility or an outcome.

The chance of at least one late line also depends on how many lines are due. Ten delivery commitments create more opportunities for a miss than one. Compare forecasts and baselines within similar delivery volumes and part or site mixes; report the eligible line count beside each supplier prediction. If your decision concerns individual orders, consider a line-level event instead of a supplier-level one.

Keep probability separate from the confidence or quality of the supporting evidence. A 70% event probability means a stated likelihood for that event within that window. It does not mean the explanation is 70% correct, or that 70% of all Acme orders will be late.

Define the decision as well. Is the team making a supplier call, asking for a recovery plan, or considering alternate supply? Those actions have different costs, so an alert threshold appropriate for a phone call may be unsuitable for an expensive sourcing change.

Replay history using only what was known at the time

Pick historical decision dates and reconstruct the evidence available on each date. If the forecast is made on August 1, an August 10 receipt and an August 15 news report cannot be inputs. Including them gives the model hindsight and makes the result look better than a real decision could have been.

Check availability time as well as event time. A disruption may have happened on July 30 but only been captured in your source on August 3. It was not available to an August 1 workflow. Revised delivery dates, corrected supplier matches and retrospectively edited records need the same care.

Freeze or retain the evidence snapshot, cutoff, published query version, model and forecast configuration version for each prediction. If historical snapshots are missing, label the exercise as a limited retrospective analysis. You can learn from it, but it is not a faithful replay of what your team would have known.

Compare against the way your team already works

Use a simple baseline alongside the proposed forecast: your current supplier watchlist, the previous period’s delivery performance, or another documented process. Apply the same event definition, decision dates and eligible subjects to both. Otherwise, the comparison can reward differences in scope rather than better decisions.

Keep model selection and final evaluation separate. You can use an earlier period to choose a prompt or threshold, then evaluate the frozen choice on a later period. Repeatedly adjusting the approach after seeing the test outcomes turns the test set into development data.

For repeated supplier forecasts, take account of overlapping windows and repeated subjects. Ten weekly alerts about the same disruption are not ten independent successful warnings. Report the number of subjects, decisions and distinct events alongside aggregate measures.

Count useful warnings and false alarms

Imagine one prediction each for 100 eligible suppliers. Twenty experience the defined late-delivery event during the window. A selected threshold flags 30 suppliers: 15 experience the event and 15 do not. Of the 70 unflagged suppliers, five experience the event and 65 do not.

Illustrative outcomes at one chosen alert threshold
DecisionEvent occurredEvent did not occurTotal
Flagged151530
Not flagged56570
Total2080100

Evaluate the probabilities, not only the alert threshold

Threshold counts tell you what happens when you turn forecasts into alerts. They do not tell you whether a stated probability is well calibrated. Across enough comparable predictions near 20%, does the event happen roughly one time in five? A few outcomes cannot establish this; show the sample size and distribution.

For a binary event, the Brier score averages the squared difference between the predicted probability and the observed outcome, encoded as 1 or 0. A probability of 0.7 for an event that occurs contributes (0.7 − 1)² = 0.09. The same probability for an event that does not occur contributes (0.7 − 0)² = 0.49. Lower average scores are better under the same evaluation setup.

A lower Brier score does not, by itself, establish better calibration. It reflects several aspects of forecast quality. Inspect reliability across probability bands as well as the aggregate score, using enough outcomes to make the comparison meaningful.

Keep unknown, ambiguous and unresolvable outcomes separate from known non-events. Report how many forecasts could be evaluated and why others could not. A good score on a selective subset can hide poor evidence coverage.

Ask whether the warning arrived early enough to help

A correct warning the day before a missed delivery may be too late to arrange alternate supply. Measure the time between the decision and the event, and compare it with the time required for the intended action. Record whether the suggested response was feasible and whether the evidence supported it.

During a pilot, log the action your team took and its timing. If a buyer intervenes and prevents a delay, the observed outcome changes. You cannot simply label every prevented event a forecasting error or claim every avoided delay as proven savings. Outcome records and intervention notes help explain the limits of the evaluation.

Review important misses and false alarms individually. Was the supplier identity wrong? Was evidence unavailable? Was the event definition unsuitable? Did the forecast ignore a contradiction? These findings are more useful for improvement than a single headline percentage.

Run a bounded evaluation in Semogram

Start with a published evidence query for one business question. Set up the forecast with its instructions, expected inputs and outputs, model, execution settings and outcome policy. Each saved forecast configuration uses a specific published query version; a later query update must be selected deliberately.

Run predictions with a stated subject, horizon and evaluation date. Inspect the saved evidence, warnings and output fields. When the outcome can be established under the policy, record what happened with observation time and supporting notes, then review the evaluation report.

Recording an outcome does not automatically retrain the model or promise a better rerun. Use what you learn to deliberately change the data, query, prompt, skill or forecaster version, and measure the new version under a comparable evaluation. Begin with a small pilot and a limited decision before expanding reliance.

  • One observable event and a written eligibility policy.
  • Historical evidence with honest availability cutoffs.
  • A current-process baseline and a held-out evaluation period.
  • Results including misses, false alarms, coverage and useful lead time.
  • Versioned changes and a prospective check on new decisions.

Apply this to your own operations

Start with one question and the records behind it. We can help you scope a workflow your team can inspect, correct and evaluate.