Your exam. Your pass mark. Your gate.
Models change, and the firms that run on them rarely know when. With GoldenEval a model works for you only after it has sat your exam: your approved test set, marked against pass marks you set, with a recorded scorecard and one-click rollback. With GoldenEval-Live the exam is re-run on schedule against every model in service, and autonomy steps back to your people the moment a result falls below the mark.
The exam.
Every model, every version, every change of assignment sits the same exam before it reaches live work: your approved test set, drawn from your own book and the baseline library, scored case by case against the pass marks your firm sets. The scorecard is written to the record. You activate the model against your criteria, or you don’t. Any change is reversible in one step, and the decision and its evidence stay in the record.
Yours stays yours.
The exam never stops.
We have no control over a hosted model, which can change beneath a version number and can vary in performance. There is also a potential responsibility to demonstrate management of the model through monitoring, calibration and documentation. A model that degraded overnight shows nothing until an error is identified, and a quarterly reporting cycle may not identify it for many weeks. GoldenEval-Live re-runs your exam on a schedule you set against each model in service, and responds on a ladder.
Model performance drifts without an outward sign. Identical questions can draw different answers between sessions; a pinned version is not a guarantee of a pinned behaviour.
Treat a model as you treat a person with approval rights: examine before, re-examine in service, withdraw the mandate on evidence — and keep the exam in your hands.
Designed against the post-market-monitoring expectations of the EU AI Act (Art 72), ISO/IEC 42001 and the NIST AI RMF — continuous monitoring feeding documented corrective action. Designed against, not certified. GoldenEval watches the models; Countercheck watches the reviewing.
Your approved models, your exams, your gate.
Model-agnostic, multi-model and model-flexible by design. Which models run, on which work, at what level of effort — and at what cost, recorded per task — is your decision, evidenced on every change.