Golden sets and iteration: evals that do not lie to you

In short

A trustworthy golden set fixes the task, source artifact, answer key, scoring rules and denominator before a run. Experts author or verify the key independently, graders remain blind to system identity, deterministic checks replace model judgment where possible, and omissions are scored separately from contradictions. Small targeted probes should precede broad sweeps.

A golden set is not merely a folder of documents with approved answers. It is a measurement instrument. If its source artifacts, answer keys or scoring rules move between runs, the resulting score cannot establish whether the system improved. It may only show that the test changed.

This matters acutely in credit-document analysis. A response can identify the correct restricted payments basket but omit the builder basket conditions. It can quote the right leverage threshold from the wrong definition. It can produce polished prose that reverses the effect of a proviso. A single pass/fail label hides these differences, even though they imply different risks for an analyst relying on the output.

The discipline is therefore less about accumulating examples than controlling what each example proves. Every reported score should answer four questions: what was measured, on which exact artifact, against whose answer key, and under what grading procedure.

What should a golden-set item contain?

The atomic unit should be a test case, not just a document. One agreement may support dozens of cases, each testing a distinct capability: retrieval, calculation, defined-term resolution, exception handling, synthesis or refusal when the text does not support an answer.

A useful test-case record contains:

FieldWhat it controls
TaskThe precise question or instruction presented to the system
Source artifactThe exact document file, version and relevant amendments
Expected answerThe substantive result supported by the artifact
EvidenceThe text and definitions that justify the expected answer
Scoring ruleThe conditions for correct, incomplete, unsupported and incorrect
Failure tagsThe capability and document feature under test
ProvenanceWho drafted, reviewed and approved the key
Set membershipCore regression, targeted probe, challenge set or exploratory set

The expected answer should specify required propositions rather than prescribe one ideal paragraph. For a covenant question, those propositions might include the applicable basket, its capacity basis, conditions to use it, interaction with other capacity and any material proviso. This permits legitimate variation in wording without treating eloquence as correctness.

Evidence is equally important. An answer key without a support trail merely relocates the trust problem. Reviewers should be able to reconstruct the conclusion from the document, including the definitions and cross-references on which it depends.

How do answer keys become contaminated?

The most obvious contamination occurs when the system being evaluated also authors the key. If a model misreads a proviso consistently, it may give the same wrong answer during key creation and evaluation. Agreement with itself then appears as success.

Model assistance can accelerate candidate generation, but the approved key needs independent verification by someone competent to interpret the artifact. The reviewer must inspect the underlying language, not simply decide whether the proposed answer sounds plausible. Material ambiguities should remain visible in the key rather than being forced into false certainty.

A more subtle problem arises when key authors have seen the candidate system’s answer. They may unconsciously shape the expected answer around its structure or excuse omissions because they understand what it was trying to say. Key creation and adjudication should therefore be separated where practical.

The provenance record need not be elaborate. It should establish who proposed the key, who checked it, which artifact they used and whether either person saw system output before approval.

Why must the denominator remain fixed?

A score is a fraction. If the denominator changes, comparison becomes unsafe.

Suppose one run excludes cases that failed because a document could not be parsed, while another includes them. The later score may fall even though answer quality improved. Conversely, removing difficult or disputed cases can create apparent progress without any change to the system.

Each run should preserve:

  • The total number of eligible cases.
  • The number attempted, omitted, failed before answering and excluded.
  • The reason for every exclusion.
  • Results on an unchanged core regression set.
  • Results on any newly added or revised cases as a separate cohort.

Do not silently rewrite a key and replace the historical result. Version the case. A corrected key may invalidate an old comparison, but that is more honest than manufacturing continuity.

Aggregate scores should also preserve relevant slices. A stable overall result can conceal a regression in amendment handling or multi-hop definition tracing if easier retrieval cases dominate the set. Slices should follow meaningful failure mechanisms, not whichever categories make the dashboard look orderly.

When should grading be deterministic?

Use deterministic checks whenever correctness can be expressed without interpretation. Model judgment introduces variance and can reward surface qualities that have little connection to the task.

Exact or programmatic checks are appropriate for document identity, currency, dates, numerical thresholds, enumerated basket names, citation presence, JSON structure and arithmetic derived from fixed inputs. Normalisation can accommodate harmless formatting differences without turning the check into an essay contest.

Some conclusions require expert judgment. A covenant summary may be accurate despite using different language from the key. In those cases, use a structured rubric based on propositions:

DimensionExample grading question
CorrectnessDoes any statement contradict the operative document?
CompletenessAre all required propositions present?
SupportCan each material claim be traced to the artifact?
ScopeDoes the answer identify assumptions and relevant limitations?
Citation qualityDoes the cited text actually support the proposition?

A model grader, if used, should receive the source evidence, approved propositions and explicit error definitions. It should not know which model produced the answer, whether the answer came from a baseline or candidate system, or what result the evaluator hopes to see. Blindness will not make a weak rubric sound, but it removes one avoidable source of bias.

Fluency should be scored only when the workflow genuinely depends on it. A well-structured answer that reverses the meaning of an exception remains wrong.

Why is omission different from error?

A missing conclusion is not equivalent to a false conclusion. Nor is an unsupported assertion equivalent to either.

Consider a question about permitted debt capacity. A response might:

  • State the correct general basket but omit a ratio condition.
  • State that the ratio condition does not apply when it does.
  • Add a capacity figure that cannot be derived from the provided artifact.
  • Decline to calculate capacity because the necessary financial inputs are absent.

These outcomes require different remediation. The first suggests incomplete extraction or synthesis. The second indicates interpretation failure. The third suggests unsupported generation. The fourth may be the correct behaviour.

A useful error taxonomy separates at least:

  • Correct and complete: All required propositions are present and supported.
  • Correct but incomplete: Present statements are supportable, but required content is missing.
  • Unsupported: A material claim lacks support in the supplied artifact or inputs.
  • Contradictory: A material claim conflicts with the operative language.
  • Unanswerable handled correctly: The system identifies missing evidence or ambiguity.
  • Operational failure: The system does not produce an assessable answer.

Teams may weight these categories differently for a particular use case. A lawyer drafting an issues list may value conservative omission differently from an analyst screening a large portfolio. The raw categories should remain available beneath any weighted score.

How should challenge cases be selected?

Large random sweeps often add volume without adding much diagnostic signal. Repeating straightforward questions across similar documents can tighten an average while leaving the consequential edge cases untouched. This is a triage heuristic about where to spend, not a claim that small samples measure better than large ones.

It does carry a cost that has to be managed explicitly. Agent runs are not reproducible: the same input can produce different tool calls and a different verdict on a second attempt. In a small probe that variance is large relative to the effect being measured, so a single flipped verdict is not evidence of anything. Before concluding a change helped or hurt, re-run the affected cases and check whether the verdict is stable. A score reported without saying how many runs it rests on is a third way for an eval to mislead, alongside a shifting denominator and a compromised key.

Start from a failure hypothesis. If a change may affect defined-term resolution, build the smallest probe containing nested definitions, alternative definitions and references through an amendment. If citation generation changed, test citations that require distinguishing operative text from a definition, schedule or superseded provision.

Challenge cases should represent mechanisms that plausibly break the system:

  • Exceptions separated from the principal rule.
  • Definitions that incorporate other definitions.
  • Amendments that replace or qualify original language.
  • Similar covenant names with different operative effects.
  • Tables, schedules or exhibits carrying essential terms.
  • Questions that cannot be answered from the provided artifact.
  • Calculations requiring both documentary terms and external inputs.

Once a probe exposes a regression, preserve a representative case in the core set. Do not allow the core set to become an archive of every bug ever observed. Redundant cases increase review cost and can make one historical failure mode dominate the aggregate result.

What does a trustworthy iteration loop look like?

Begin with a narrow change and a narrow prediction. State which cases should improve, which should remain unchanged and what new failure could emerge. Run the targeted probe first. If it fails, a broad sweep rarely adds useful information.

When the probe passes, run the fixed core set and inspect category-level changes. Review raw outputs for all changed verdicts, not only the aggregate score. A grader change deserves the same scrutiny as a system change because it alters the measurement instrument.

Newly discovered failures should enter a quarantine or candidate set until their keys are independently reviewed. Only then should selected cases join the fixed regression set. This prevents hurried debugging examples from quietly weakening the benchmark.

The resulting evaluation record should let another reviewer reconstruct the claim: system version, prompt and tool configuration; artifact versions; case-set version; answer-key provenance; grader version; exclusions; category counts; and adjudicated disagreements.

CreditGPT or any other document-analysis system can be assessed under this structure. The product name matters less than the controls. A defensible eval does not ask whether an output looks good. It establishes which proposition was tested, what evidence determined the answer and whether the same measurement can be repeated after the next change.

Common questions

How large should a golden set be for an AI document-analysis system?

Start with the smallest set that can expose the failure mode under investigation. Expand only when additional examples represent distinct document structures, drafting variants or economically important edge cases, rather than duplicating questions the set already tests.

Can another language model create the answer key for an AI eval?

A model can propose candidate answers, but those answers should not become the key without independent review against the source document. Using substantially the same system to author and answer the test can preserve its blind spots and make agreement look like correctness.

Should missing information and incorrect information receive the same eval score?

No. An omission, an unsupported assertion and a contradiction of the document create different risks and usually require different fixes. Record them separately even if a downstream roll-up applies weights for a particular workflow.

Related

See this run against your own documents.

Book a demo