Applied AI insights
Evaluation and adoption
What is a golden evaluation set?
A golden evaluation set is a versioned collection of representative cases from a real workflow, each paired with an agreed expected outcome. Teams use it to test whether an AI system meets task-specific acceptance thresholds before launch and after changes to its model, prompts, tools or source data.
Public benchmark: How capable is the model?
Golden evaluation set: Is this system reliable enough for this workflow?
Terminology varies
Teams may use terms such as golden dataset, evaluation set, test set or benchmark set differently. In this guide, a golden evaluation set means representative workflow cases with agreed expected outcomes, version control and task-specific acceptance thresholds.
Why generic benchmarks are not enough
Public benchmarks measure a model’s general ability on someone else’s tasks. Your invoices, manifests and product sheets carry their own formats, abbreviations, edge cases and failure modes. A model can score impressively in public and still misread the one field your billing depends on. There is no general score that certifies a system for your workflow; only a test built from your workflow can do that.
What belongs in a golden evaluation set
A useful set is more than a folder of examples. It should contain:
- Common, representative cases
- Legitimate variations in format and language
- Known failure cases
- High-consequence edge cases
- Expected outcomes confirmed by people who understand the work
- The metadata needed to evaluate each case
- Acceptance thresholds
- Version information
- A record of changes and disagreements
When two subject-matter experts disagree about the expected outcome of a case, that disagreement is a useful finding about the workflow itself. It must be resolved and recorded, not discarded as noise.
How to build one
- 01Sample the real distribution. Pull cases from actual history: the clean majority, the messy scans, the supplier who formats everything differently, the month-end spikes.
- 02Record the correct answer for each. The people who do the work today confirm the expected outcome, field by field. Disagreements between experts are findings, not noise.
- 03Include known failure cases. Duplicates, missing pages, wrong-currency invoices, renamed SKUs: the cases that burnt you before go in deliberately.
- 04Agree acceptance thresholds before testing. Decide per field what accuracy is acceptable and which errors are never acceptable, so the go / no-go decision is made in advance, not negotiated after the results arrive.
- 05Keep it versioned and rerun it. Every change to models, prompts or source data reruns the set. Yesterday’s pass does not carry over; regressions are caught before users find them.
How large should it be?
There is no universal number. The set should be large and varied enough to represent normal work, material edge cases and high-consequence failures. It should then grow deliberately as production reveals new cases.
The coverage a workflow needs depends on:
- How much the workflow varies day to day
- The consequences of an error
- The number of fields or decisions being evaluated
- The representative data actually available
- How often rare but serious cases occur
- The confidence the release decision requires
Worked example
Take the supplier inbox to ERP workflow: supplier emails and attachments are read, approved fields extracted, records matched, and exceptions routed for approval. A slice of its golden set might look like this.
| Representative case | Expected outcome | Failure being tested | Acceptance decision |
|---|---|---|---|
| Clean supplier invoice | All approved fields extracted and matched to the purchase order | Baseline extraction accuracy | Must pass across the sample |
| Poor-quality scan | Uncertain fields flagged for human review, none guessed | Silent misreads on low-quality input | Always flagged, never guessed |
| Duplicate invoice | Flagged as a duplicate of the earlier record | Double processing and double payment | Never passes silently |
| Missing purchase-order number | Routed to the exception queue with the reason recorded | Unmatched postings | Must route, never post |
| Wrong currency | Currency mismatch raised before any posting | Cross-currency posting errors | Never passes silently |
| Renamed SKU | Linked to the current item, or escalated when unsure | Stale product mappings | Escalate below agreed confidence |
| Conflicting totals | Discrepancy raised with both sources referenced | Reconciliation errors | Must raise, never average |
| Unauthorised supplier | Blocked and escalated to the accountable owner | Payments outside policy | Never passes silently |
Each row pairs a real case type with the outcome the team has agreed is correct and the decision rule that applies to it. The thresholds themselves are set per engagement — the same principle as agreeing the measure before the build.
How to use it before launch and after changes
The same versioned set is the release gate before launch and the regression test afterwards. Rerun it whenever:
- The model changes
- Prompts change
- Tools or integrations change
- Source documents or data formats change
- Business rules change
- A production failure reveals a new case — which then joins the set
A previous pass does not carry over after a material system change. The set is only evidence about the system that was tested; when the system changes, the evidence must be regenerated. In an AgentAligned engagement this rerun discipline is part of Sustaining Production, alongside monitoring and drift investigation.
What it does not do
A golden set cannot promise zero errors; no responsible team promises that. It bounds the risk: it tells you how the system behaves on a representative slice, so human review can be placed where the consequence of an error is high. It also does not replace production monitoring — live data drifts, and the set itself must grow as new edge cases appear.
In an AgentAligned engagement, the golden set is built in the Bottleneck Diagnostic and the acceptance thresholds are agreed before the build begins. Where cases contain personal information, the set is handled under the same POPIA-aligned controls as the workflow itself: minimised, access-controlled and kept in an approved environment.
Sources and further reading
This article describes AgentAligned’s practical approach. For broader context on measuring and governing AI systems:
- NIST AI Risk Management Framework (NIST) — a widely used framework for identifying, measuring and managing AI risk; its Measure function covers evaluation practice.
- OpenAI Evals (GitHub) — an open-source example of tooling for defining and running evaluation sets against language-model systems.
- Protection of Personal Information Act 4 of 2013 (justice.gov.za) — relevant where evaluation cases contain personal information; see our POPIA and sovereign AI guidance.