GoldenEval · a Countersign platform capability

Your exam. Your pass mark. Your gate.

Models change, and the firms that run on them rarely know when. With GoldenEval a model works for you only after it has sat your exam: your approved test set, marked against pass marks you set, with a recorded scorecard and one-click rollback. With GoldenEval-Live the exam is re-run on schedule against every model in service, and autonomy steps back to your people the moment a result falls below the mark.

GoldenEval is your tool for model management in a time of rapid change — fine-tuning which models are assigned to which processes and risk ratings, based on your priorities of cost or performance. There is considerable opportunity for model engineering and model management across the service.
Before a model works for you

The exam.

Every model, every version, every change of assignment sits the same exam before it reaches live work: your approved test set, drawn from your own book and the baseline library, scored case by case against the pass marks your firm sets. The scorecard is written to the record. You activate the model against your criteria, or you don’t. Any change is reversible in one step, and the decision and its evidence stay in the record.

APPROVEOnly models your firm approves are available. Versions are locked; nothing updates beneath you.
ASSIGNA model and an effort level per task and risk class — stronger models on higher-risk work, efficient models on routine internal steps.
EXAMINEThe assignment sits your test set. Each case scored. Pass marks yours.
ACTIVATEYou activate against your criteria. The scorecard and the decision are recorded.
ROLL BACKOne click returns the previous model. Recorded, like everything else.
Three test libraries

Yours stays yours.

Baseline libraryThe exam every deployment inherits — the common fund-administration cases.
Your private libraryTests your team adds from your own book; recurring issues become permanent cases. Held in your tenancy, never shared.
Shared library opt-inOff by default. Share your test library with other Countersign clients — for example in the interests of the industry for payments processing. Client data never enters this test suite.
GoldenEval-Live · in service

The exam never stops.

We have no control over a hosted model, which can change beneath a version number and can vary in performance. There is also a potential responsibility to demonstrate management of the model through monitoring, calibration and documentation. A model that degraded overnight shows nothing until an error is identified, and a quarterly reporting cycle may not identify it for many weeks. GoldenEval-Live re-runs your exam on a schedule you set against each model in service, and responds on a ladder.

RE-RUNYour exam, on schedule, against each live model.
ALERTA result below the pass mark raises an exception to your administrator first.
SAMPLE MOREStaff review sampling rises on the affected work.
STEP BACKAny policy-based release on that work suspends itself; items return to per-item staff release.
ROLL BACKOne click to the last model that passed. Every step recorded.
↩ a model that passes again can be re-activated — by you
Why it exists
The known failure

Model performance drifts without an outward sign. Identical questions can draw different answers between sessions; a pinned version is not a guarantee of a pinned behaviour.

The design answer

Treat a model as you treat a person with approval rights: examine before, re-examine in service, withdraw the mandate on evidence — and keep the exam in your hands.

The standard we hold

Designed against the post-market-monitoring expectations of the EU AI Act (Art 72), ISO/IEC 42001 and the NIST AI RMF — continuous monitoring feeding documented corrective action. Designed against, not certified. GoldenEval watches the models; Countercheck watches the reviewing.

Scoring methods, schedules and thresholds are configured by your firm and described in the technical pack, under NDA. Any statement about AI Act status is Countersign’s own assessment, not a determination by any authority.

Your approved models, your exams, your gate.

Model-agnostic, multi-model and model-flexible by design. Which models run, on which work, at what level of effort — and at what cost, recorded per task — is your decision, evidenced on every change.