The benchmark behind the numbers
Every accuracy figure on certusqa.com comes from one frozen, hand-labelled suite run against the judge that classifies test failures. This page publishes the whole result, misses included, so you do not have to take the headline on trust.
Per-class precision and recall
Support is the number of hand-labelled cases in each class. The two classes marked with a gate floor must reach recall 1.0 or the evaluation fails: a missed real bug or a guessed pass is the error we treat as unacceptable.
| Class | Support | Predicted | Precision | Recall | F1 |
|---|
The misses
Every case the judge got wrong, by class. Case bodies are not published; the direction of each miss is.
| Artifact | Human label | Judge said | Confidence |
|---|
Run history
Rows with a different suite fingerprint are the earlier 12-case suite and are not comparable to the 37-case runs.
| Evaluated | Cases | Accuracy | Macro F1 | Judge | Suite |
|---|
Read this honestly
This page renders golden-eval.json. If the tables above are empty, open the file directly; it is the same data.