Evaluation Plan
Nothing on this page is a result. These are the evaluations designed to test claims that are not yet supported, with their pass criteria fixed in advance so that a weak outcome cannot be reinterpreted as a success.
What can already be argued
Some advantages follow from the structure itself and do not need an experiment:
- Compared with a manual spreadsheet, source, transformation, and judgment grounds are structurally linked, which makes re-review and change-impact tracing tractable.
- Compared with a pipeline that reports only successes, unprocessed, unsupported, and review-required states are visible, so omissions can be found at all.
- Compared with a bare file hash, a revision's change reason and current-use state can be explained.
- A result commitment and revision state can, by design, be verified without publishing source text on-chain.
These are statements about structure. They are not claims that a user finds omissions faster, which is an empirical question.
What requires evidence
| Evaluation question | Method | Pre-registered pass criteria |
|---|---|---|
| Are omissions and duplicates easier to find? | Review the same fixture through existing manual/integration output and through the coverage interface | Measure detection rate against a known answer set, time to detection, and false positives |
| Is there no silent drop? | Exhaustive comparison of terminal outcomes across supported, unsupported, and failing fixtures | Every in-scope input has either a result or an explicit state |
| Can a result be explained? | User task test tracing a result back to its source | 4 of 5 participants locate the source, policy, and exception |
| Does review productivity improve? | Cross-review by 3 tax professionals of existing materials versus an Evidence Pack | 2 or more report ≥30% reduction in preparation and review time, with zero critical omissions |
| Is it reproducible? | Repeated runs on identical input and version, plus revision diff verification | Zero change in canonical IDs; where versions differ, the cause is explainable |
| Does GIWA verification work? | Valid, tampered, and superseded revision scenarios | Valid current revision passes; tampered and superseded revisions are rejected |
Why the criteria are fixed in advance
Each row commits to a threshold before the measurement exists. This is deliberate: the same data can support a favorable reading after the fact, and a comparison whose success condition is decided afterwards is not evidence.
Two of these evaluations depend on interfaces that do not exist yet — the coverage interface and the Evidence Pack. They cannot be run until Tax Inventory & Report is built. The GIWA scenarios cannot be run until a contract is deployed; see GIWA Integrity Verification.
Relationship to current claims
Until these evaluations produce results, this documentation does not use comparative language about speed, accuracy, or omission rates. See Claim Boundaries.