Applied AI insights

Evaluation and adoption

What is a golden evaluation set?

A golden evaluation set is a versioned collection of representative cases from a real workflow, each paired with an agreed expected outcome. Teams use it to test whether an AI system meets task-specific acceptance thresholds before launch and after changes to its model, prompts, tools or source data.

Public benchmark: How capable is the model?

Golden evaluation set: Is this system reliable enough for this workflow?

Terminology varies

Teams may use terms such as golden dataset, evaluation set, test set or benchmark set differently. In this guide, a golden evaluation set means representative workflow cases with agreed expected outcomes, version control and task-specific acceptance thresholds.

Why generic benchmarks are not enough

Public benchmarks measure a model’s general ability on someone else’s tasks. Your invoices, manifests and product sheets carry their own formats, abbreviations, edge cases and failure modes. A model can score impressively in public and still misread the one field your billing depends on. There is no general score that certifies a system for your workflow; only a test built from your workflow can do that.

What belongs in a golden evaluation set

A useful set is more than a folder of examples. It should contain:

  • Common, representative cases
  • Legitimate variations in format and language
  • Known failure cases
  • High-consequence edge cases
  • Expected outcomes confirmed by people who understand the work
  • The metadata needed to evaluate each case
  • Acceptance thresholds
  • Version information
  • A record of changes and disagreements

When two subject-matter experts disagree about the expected outcome of a case, that disagreement is a useful finding about the workflow itself. It must be resolved and recorded, not discarded as noise.

How to build one

  1. 01Sample the real distribution. Pull cases from actual history: the clean majority, the messy scans, the supplier who formats everything differently, the month-end spikes.
  2. 02Record the correct answer for each. The people who do the work today confirm the expected outcome, field by field. Disagreements between experts are findings, not noise.
  3. 03Include known failure cases. Duplicates, missing pages, wrong-currency invoices, renamed SKUs: the cases that burnt you before go in deliberately.
  4. 04Agree acceptance thresholds before testing. Decide per field what accuracy is acceptable and which errors are never acceptable, so the go / no-go decision is made in advance, not negotiated after the results arrive.
  5. 05Keep it versioned and rerun it. Every change to models, prompts or source data reruns the set. Yesterday’s pass does not carry over; regressions are caught before users find them.

How large should it be?

There is no universal number. The set should be large and varied enough to represent normal work, material edge cases and high-consequence failures. It should then grow deliberately as production reveals new cases.

The coverage a workflow needs depends on:

  • How much the workflow varies day to day
  • The consequences of an error
  • The number of fields or decisions being evaluated
  • The representative data actually available
  • How often rare but serious cases occur
  • The confidence the release decision requires

Worked example

Take the supplier inbox to ERP workflow: supplier emails and attachments are read, approved fields extracted, records matched, and exceptions routed for approval. A slice of its golden set might look like this.

Illustrative example, not a reported client result.
Example golden evaluation set cases for a supplier inbox to ERP workflow
Representative case Expected outcome Failure being tested Acceptance decision
Clean supplier invoiceAll approved fields extracted and matched to the purchase orderBaseline extraction accuracyMust pass across the sample
Poor-quality scanUncertain fields flagged for human review, none guessedSilent misreads on low-quality inputAlways flagged, never guessed
Duplicate invoiceFlagged as a duplicate of the earlier recordDouble processing and double paymentNever passes silently
Missing purchase-order numberRouted to the exception queue with the reason recordedUnmatched postingsMust route, never post
Wrong currencyCurrency mismatch raised before any postingCross-currency posting errorsNever passes silently
Renamed SKULinked to the current item, or escalated when unsureStale product mappingsEscalate below agreed confidence
Conflicting totalsDiscrepancy raised with both sources referencedReconciliation errorsMust raise, never average
Unauthorised supplierBlocked and escalated to the accountable ownerPayments outside policyNever passes silently

Each row pairs a real case type with the outcome the team has agreed is correct and the decision rule that applies to it. The thresholds themselves are set per engagement — the same principle as agreeing the measure before the build.

How to use it before launch and after changes

The same versioned set is the release gate before launch and the regression test afterwards. Rerun it whenever:

  • The model changes
  • Prompts change
  • Tools or integrations change
  • Source documents or data formats change
  • Business rules change
  • A production failure reveals a new case — which then joins the set

A previous pass does not carry over after a material system change. The set is only evidence about the system that was tested; when the system changes, the evidence must be regenerated. In an AgentAligned engagement this rerun discipline is part of Sustaining Production, alongside monitoring and drift investigation.

What it does not do

A golden set cannot promise zero errors; no responsible team promises that. It bounds the risk: it tells you how the system behaves on a representative slice, so human review can be placed where the consequence of an error is high. It also does not replace production monitoring — live data drifts, and the set itself must grow as new edge cases appear.

In an AgentAligned engagement, the golden set is built in the Bottleneck Diagnostic and the acceptance thresholds are agreed before the build begins. Where cases contain personal information, the set is handled under the same POPIA-aligned controls as the workflow itself: minimised, access-controlled and kept in an approved environment.

Sources and further reading

This article describes AgentAligned’s practical approach. For broader context on measuring and governing AI systems:

About the author

The AgentAligned Editorial Team writes from the delivery side of applied AI engagements: the practitioners who design, build and run governed workflows for South African operations. Articles are reviewed before publication and updated when practice changes. More about how we work.