Evals · pipeline v0.2
These accuracy numbers are not a quality claim
0 of 5 label files are marked verified. The rest were drafted by the same model being graded, so the accuracy below measures the model agreeing with itself — which is why it reads 100.0%. It is a working harness, not a result.
Verifying the labels is the one task blocking every measurable claim in this project. See corpus/labels/README.md for the procedure and the conventions to settle first.
Partial matches score half. Uniform bars are the tell: every field scores identically because each label was drafted from the extraction it is grading.
The question that decides whether confidence scoring earns its keep: of the fields sent to a human, how many were actually wrong — and of the wrong fields, how many did routing catch?
Both read 0% because the labels report zero errors to catch — the tautology again. This number only becomes meaningful once the labels are verified and real disagreements appear.
The gate on everything above.
Flagged during drafting — each needs a human decision before its label can be marked verified.
| Document | Accuracy | Cost | Latency | Labels |
|---|---|---|---|---|
| acm-research-oregon-lease | 100.0% | $0.14 | 20s | verified |
| dexcom-office-lease2 in review | 100.0% | $0.45 | 1m 8s | verified |
| hyliion-industrial-lease6 in review | 100.0% | $0.04 | 27s | verified |
| tenaya-lab-lease3 in review | 100.0% | $0.12 | 20s | verified |
| yoshiharu-retail-lease6 in review | 100.0% | $0.18 | 25s | verified |