03 · How a support claim is built
GitHub issues six kinds of token. Which ones do you catch?
This structure exists to answer that honestly.
The problem
A detector name is the wrong unit
The detector called github-token matches ghp_ | gho_ | ghu_ | ghs_ | ghr_, five shapes. GitHub issues six. For a long time nothing covered the sixth, and no rule list would have told you.
A support claim counts credentials a provider actually issues, not detectors we happened to write.
One file records that list, whatever the code happens to catch.
The three inputs
Three files, three different questions
They are kept apart on purpose: each one can be wrong on its own, and each is written by a different process.
- Written by
- A generator, from the core's detector registry.
- Pinned to
- One exact commit: the build the measurement ran against.
- Holds
- A flat list of detector ids and titles, with nothing about quality.
- Written by
- People, by hand, with a source link or a written note for every entry.
- Pinned to
- Nothing. It describes providers, not our code.
- Holds
- Every credential family, whether or not we detect it.
- Written by
- A run of the benchmark corpus.
- Pinned to
- A manifest of corpus hashes plus the version measured.
- Holds
- A verdict per detector: status, evidence tier, and why.
The join
Measurement lands on detectors. Claims land on families.
docs/support-matrix.md and the README tableBoth generated. npm run ci fails if either drifts from the JSON.The one-way direction matters: the matrix reads the taxonomy, and never the reverse. A code review rule says so outright: variant support must not be inferred from a related family name. Sharing a prefix with a supported token is not evidence.
The clever part
A family with no detector is still a row
In the taxonomy, a credential nobody wrote a detector for has an empty list: "detectors": []. That empty list produces an unsupported row in the published table, and it doubles as a worklist. GitHub's sixth token started here.
Usual
Ship a rule list. A credential nobody wrote a rule for simply is not mentioned. You cannot tell “we checked and do not catch it” apart from “we never thought about it”.
Here
Lists the gaps in the same table as the wins, each with its reason. 25 families have no detector today, and every one of them is published.
A test enforces this: a zero-detector family must carry a source link or a written note. An unsupported claim with no reason fails the build; the spec calls that a bug in the taxonomy, not a fact about the provider.
// a real unsupported row
"provider": "atlassian",
"status": "unsupported",
"detectors": [],
"reason": "the pinned third-party detector targets the distinct ATCT
access-token family, not the ATAT-prefixed API-token shape."Not every lookalike counts
Some things are left out on purpose
If anything key-shaped became a family, unsupported would fill with things that were never secrets. So a few are excluded by name, with the reason recorded.
- AWS AIDA…
- Twilio Account SID
- Twilio API Key SID
- Stripe pk_…
- Supabase sb_publishable_…
Identifiers, not secrets
An AWS AIDA value names an IAM user. A Twilio SID identifies an account: it gates detection of the token beside it, but leaking it is not a leak.
Documented as public
Stripe's publishable key and Supabase's publishable key are both meant to ship in browser code. Their own vendors say so.
As a result, unsupported always means “a real credential this project does not catch”. It never means “a string that resembles our patterns”.
What a row says
A provisional row shows its arithmetic
When a family falls short, the row carries the exact gates it missed, with the numbers.
// a real provisional row
"provider": "anthropic",
"status": "provisional", "evidenceTier": "T1",
"evidenceBasis": "provider-documented",
"reason":
documented.minimumPositiveCases: 5 < 6
documented.minimumPositiveAxes: 2 < 4
documented.minimumBenignCases: 5 < 8
documented.minimumControlAxes: 3 < 4That family's token format is provider-documented, the best evidence tier there is. It was still provisional because its fixtures were one positive case and three benign controls short. Nobody can argue with that; somebody can go and write four fixtures. Most of the matrix has since been filled in that way.
Status distribution
The shipped matrix, 108 rows
- Stable 83 — clears the bar; safe to rely on
- Unsupported 17 — published anyway, with the reason
- Provisional 7 — useful, but the evidence is incomplete
- Pending 1 — neither a detection nor a miss can be trusted yet
The 83 stable rows are counted in two separate streams: 57 rest on a contract the provider published, and 26 were earned by measurement alone because no vendor documentation exists. Evidence tiers are graded separately (T1 57, T2 29, T3 4, T0 1). Qualifying empirically never promotes T2 evidence to T1.
Counting carefully
Two numbers that disagree mean the pin is working
Families across 63 providers. The benchmarks repository keeps moving.
A frozen copy. It changes only when someone re-runs the measurement and re-pins.
One release ago they matched exactly. A gap between them is not an error. It shows how old the pinned evidence is. This page shows the two side by side instead of merging them.
The same goes for versions. The shipped matrix was measured on 0.1.0-beta.7 (benchmarks at cfaeac4); the current release is 0.1.0-beta.10, and its drift gate ran against exactly this matrix. Both come from the product's site feed at e748d47, read on 2026-09-28.