GAU
Gauge
Evaluation harness
What it does
Gauge is the measurement plane. Its purpose is that changing a model, a prompt or a threshold produces a number comparable to the last number rather than an impression, which needs a fixed set of cases and one scoring method.
It sits under the pipeline rather than inside it: it measures the other modules and never touches a production input.
What it accepts
System outputs compared against a reference set to track consistency and drift over time.
- The evaluation corpus
- A fixed set of cases and their expected outcomes, versioned under the same change control as the code so that changing the corpus is itself a recorded change.
- Candidate configurations
- Any model, prompt, rule set or threshold change proposed for release.
What it emits
- Scores
- The same scoring method applied to every candidate, so two runs are comparable by construction.
How it is configured
Configuration belongs to the client. It encodes what your specialists already know, and it is versioned like code.
- Corpus composition
- What the corpus contains and what each case tests.
- Scoring method
- How outcomes are scored, versioned with the corpus.
- Re-evaluation thresholds
- The score movement that triggers re-evaluation, agreed before build.
What it produces as evidence
- Evaluation records
- A versioned record of every run: the corpus version, the candidate configuration, the scores and the comparison against the previous release.
What it is tested against
Gauge is tested against acceptance criteria agreed before build: the corpus versioned and frozen, every release carrying an evaluation record, and no change shipping without a score. Gauge is also how the other modules prove their own criteria are still met after a change.