Countercheck · a Countersign platform capability

Staff review consistently monitored for quality.

Staff review is how it is done today, and that continues with the introduction of AI. With Countersign the staff review is directed based on risk assessments; with Countercheck the quality of that staff review is measured and monitored. AI platforms typically include a person’s review — but can they show it still means anything six months in, when volumes are high, drafts are usually right, and attention is fading? Countersign treats review quality as a system property: calibrated, measured, and evidenced.

A signature is only worth the review behind it. Countercheck measures whether that review is working.
How it works — the vigilance loop
CALIBRATEReview is calibrated on scored cases from your own book before approval rights go live. Signed off at onboarding.
ROUTERisk ratings — the administrator's own rules — decide what needs senior eyes, second checkers, and deliberate friction.
MEASURECatch rates, review time and outcomes read continuously from the record — at team level by default.
PROBEThe system periodically tests whether errors still get caught. Probes are built never to leave your estate.
EVIDENCECalibration, checks and outcomes land in the append-only record — a control your auditor can sample.
↩ decline in any measure triggers recalibration
Why it exists
The known failure

Research on human–AI oversight is consistent: when AI is usually right, people stop checking. There is a temptation to easy-click approve when there isn’t the immediacy of line-manager review.

The design answer

Don't assume vigilance — allocate it. Risk-tiered queues, capped volumes, friction where stakes are high, and continuous measurement of whether review is still working.

The standard we hold

Release is governed by your risk framework, enforced in the database. Countercheck completes the control: the reviewing half is measured with the same discipline as the model half — which is GoldenEval.

Probe rates, telemetry and calibration mechanics are part of the risk assessment for Service Delivery Management, and are not published.

Measured, not assumed.

Countercheck measures the review process, never individual performance as a product purpose: team-level by default, individual data role-gated, no league tables. GoldenEval watches the models; Countercheck watches the reviewing.