Skip to main content

Evaluation Plan

Planned

Nothing on this page is a result. These are the evaluations designed to test claims that are not yet supported, with their pass criteria fixed in advance so that a weak outcome cannot be reinterpreted as a success.

What can already be argued

Some advantages follow from the structure itself and do not need an experiment:

  • Compared with a manual spreadsheet, source, transformation, and judgment grounds are structurally linked, which makes re-review and change-impact tracing tractable.
  • Compared with a pipeline that reports only successes, unprocessed, unsupported, and review-required states are visible, so omissions can be found at all.
  • Compared with a bare file hash, a revision's change reason and current-use state can be explained.
  • A result commitment and revision state can, by design, be verified without publishing source text on-chain.

These are statements about structure. They are not claims that a user finds omissions faster, which is an empirical question.

What requires evidence

Evaluation questionMethodPre-registered pass criteria
Are omissions and duplicates easier to find?Review the same fixture through existing manual/integration output and through the coverage interfaceMeasure detection rate against a known answer set, time to detection, and false positives
Is there no silent drop?Exhaustive comparison of terminal outcomes across supported, unsupported, and failing fixturesEvery in-scope input has either a result or an explicit state
Can a result be explained?User task test tracing a result back to its source4 of 5 participants locate the source, policy, and exception
Does review productivity improve?Cross-review by 3 tax professionals of existing materials versus an Evidence Pack2 or more report ≥30% reduction in preparation and review time, with zero critical omissions
Is it reproducible?Repeated runs on identical input and version, plus revision diff verificationZero change in canonical IDs; where versions differ, the cause is explainable
Does GIWA verification work?Valid, tampered, and superseded revision scenariosValid current revision passes; tampered and superseded revisions are rejected

Why the criteria are fixed in advance

Each row commits to a threshold before the measurement exists. This is deliberate: the same data can support a favorable reading after the fact, and a comparison whose success condition is decided afterwards is not evidence.

Two of these evaluations depend on interfaces that do not exist yet — the coverage interface and the Evidence Pack. They cannot be run until Tax Inventory & Report is built. The GIWA scenarios cannot be run until a contract is deployed; see GIWA Integrity Verification.

Relationship to current claims

Until these evaluations produce results, this documentation does not use comparative language about speed, accuracy, or omission rates. See Claim Boundaries.