Evidence

The benchmark behind the numbers

Every accuracy figure on certusqa.com comes from one frozen, hand-labelled suite run against the judge that classifies test failures. This page publishes the whole result, misses included, so you do not have to take the headline on trust.

Per-class precision and recall

Support is the number of hand-labelled cases in each class. The two classes marked with a gate floor must reach recall 1.0 or the evaluation fails: a missed real bug or a guessed pass is the error we treat as unacceptable.

ClassSupportPredictedPrecisionRecallF1

The misses

Every case the judge got wrong, by class. Case bodies are not published; the direction of each miss is.

ArtifactHuman labelJudge saidConfidence

Run history

Rows with a different suite fingerprint are the earlier 12-case suite and are not comparable to the 37-case runs.

EvaluatedCasesAccuracyMacro F1JudgeSuite

Read this honestly

This page renders golden-eval.json. If the tables above are empty, open the file directly; it is the same data.