Skip to content

GAU

Gauge

Evaluation harness

1.0Purpose

What it does

Gauge is the measurement plane. Its purpose is that changing a model, a prompt or a threshold produces a number comparable to the last number rather than an impression, which needs a fixed set of cases and one scoring method.

It sits under the pipeline rather than inside it: it measures the other modules and never touches a production input.

2.0Input

What it accepts

System outputs compared against a reference set to track consistency and drift over time.

The evaluation corpus
A fixed set of cases and their expected outcomes, versioned under the same change control as the code so that changing the corpus is itself a recorded change.
Candidate configurations
Any model, prompt, rule set or threshold change proposed for release.
3.0Output

What it emits

Scores
The same scoring method applied to every candidate, so two runs are comparable by construction.
4.0Config

How it is configured

Configuration belongs to the client. It encodes what your specialists already know, and it is versioned like code.

Corpus composition
What the corpus contains and what each case tests.
Scoring method
How outcomes are scored, versioned with the corpus.
Re-evaluation thresholds
The score movement that triggers re-evaluation, agreed before build.
5.0Evidence

What it produces as evidence

Evaluation records
A versioned record of every run: the corpus version, the candidate configuration, the scores and the comparison against the previous release.
6.0Acceptance

What it is tested against

Gauge is tested against acceptance criteria agreed before build: the corpus versioned and frozen, every release carrying an evaluation record, and no change shipping without a score. Gauge is also how the other modules prove their own criteria are still met after a change.