04 · Evaluation methods
The scanner never gets a vote on whether it was right
There are ten test methods, and one rule holds for all of them.
The rule
Write down the answer first. Then run it.
Every method insists the expected result is authored, reviewed and hashed before any scanner touches the bytes. Without that order, the numbers mean nothing.
- Spec 01
- “A scanner result must never create or revise ground truth.”
- Spec 04
- “…rather than inferring validity from scanner agreement.”
- Spec 05
- “Recalculate spans from construction, never by searching scanner output.”
- Spec 06
- “Never relabel a control because a competitor flags it.”
- Spec 08
- “Majority agreement cannot promote, demote, or rewrite an expectation.”
Scanners never see the answer: adapters get only the fixture's identity and its bytes, never the expectation, never the evidence tier.
How a result is graded
Five answers, not two
“Did it find the secret?” is too blunt. The true span is known to the byte, so the grade asks which bytes.
- ExactPass
- CoveredPass
- OverbroadFail
- PartialFail
- MissFail
- secret bytes, redacted
- secret bytes that leaked
- innocent bytes taken too
Catching 31 of 32 bytes is a leak, not a near miss.
That is Partial, and it fails. So does Overbroad: swallowing the paragraph around the key hides the key but ruins the log line. Only a span inside its authored allowance passes.
The ten methods
Each one looks for a different way to be wrong
They are numbered because they run in order: everything after the first is derived from it. You cannot write a twin without a canonical positive to twin.
First, make a truth
one credential, one simplest case
Canonical positive
Does it find the easy one?
The seed every other method grows from. Byte ranges, evidence and source hash are recorded before a scanner runs. Passing this alone earns no support claim.
Then check it only takes what it should
a detector that matches more text looks better on positives alone
Negative twin
Change exactly one thing. Does it go quiet?
ACME_KEY=… must fire
ACMX_KEY=… must notOne mutation only: prefix, length, alphabet, boundary or public-prefix. Two changes is not a twin. The pair is scored as one unit, so you cannot win by flagging everything.
Benign lookalikes
Does it stay silent on things that merely look like keys?
AKIAIOSFODNN7EXAMPLE from the vendor's own docs
<your-api-key-here> placeholder
pk_live_… publishable, meant to be publicEach is deliberately close to the real thing and probes one specific way to overmatch. Policy-based controls are counted in a separate denominator from source-backed ones, so a judgement call cannot dilute a fact.
Then attack it
edges, characters, containers, and machine-made variants
Boundary cases
What happens one character either side of the rule?
Every documented length and delimiter needs a case inside and outside it, or a written reason why that pair is meaningless. This catches off-by-one ranges and accidental substring matches, down to end of file, a missing final newline, CRLF, and a key touching the next token.
Alphabet mutations
Is it really checking the body, or just the prefix and the length?
ACME_a9f3k2… legal characters, fire
ACME_a9f!k2… illegal character, silentOne character swapped, everything else held still. A cheap detector that only checks “right prefix, roughly right length” fails here and nowhere else.
Context permutations
Same key, different wrapper. Same answer?
bare · .env · JSON · YAML · TOML · source code · Markdown · URI · quoted · CRLF. The credential bytes do not change; everything around them does. The expected answer must survive every supported wrapper. Each context is reported separately, so strong results in one cannot hide a weak one.
Generated mutations
What about the thousand variants nobody would hand-write?
The spec says this is not fuzzing. Variants are generated before any scan, each tagged in advance as preserve, invalidate or review-required, and bounded in count and cost. With operator v3, seed 41 and the source hash, a clean checkout reproduces them byte for byte.
Then look outward, and forward
other tools, unseen cases, and the passage of time
Competitor disagreement
Where do we and Gitleaks and TruffleHog differ on identical bytes?
Competitors are observations, never an oracle. Every adapter is graded against the authored expectation on its own. A disagreement opens a review item for a human; it never edits the expectation.
Holdout evaluation
How does it do on cases it was never tuned against?
This is the exam nobody studied for. The ordinary development command cannot even read the holdout corpus. The candidate is frozen before the cases are revealed, and touching it afterwards voids the run. A newer rule closes the last loophole: holdout results may never be used to pick scorer thresholds or weights, and retuning because of what holdout showed contaminates the epoch.
Regression freeze
Could we quietly break this next month?
A finished run becomes an immutable comparison point (baselines/0.1.0-beta.9.json: fixture hashes, versions, every outcome). A new Partial, Miss or false alarm fails the build. An intentional change needs a written, approved reason before the baseline moves, and improvements never erase the old record.
A third bucket
There is a third answer
Real formats are ambiguous at the edges. A suite that forces every case into pass or fail starts inventing facts. So these methods keep a third answer.
Preserve
Still a secret after the change. Must be caught.
Invalidate
No longer a secret. Must be ignored.
Review required
Nobody knows yet. Scored as nothing; a human decides.
A review-required case makes no pass or fail claim and never enters an average. The same goes for T0 evidence, which can be observed and inspected but is deliberately left unscored.
What it all buys
Even with all ten green
Proven
That on these exact bytes, at this exact version, with these pinned tool versions, the scanner behaved as reviewers said it should, and anyone can reproduce that from a clean checkout.
Not proven
Production accuracy. Method 01 says so outright: this is fixture-relative coverage. Method 09 adds that holdout does not establish production accuracy either. Ten green methods describe the corpus, not real-world traffic.
No single method is enough on its own. That is also why the support matrix grades a family on how many of these it has cleared rather than on whether it was ever detected once.