02 · How it detects
Every detector answers the same question: how sure, and why?
There are exactly five kinds of why, and they are ranked. Most of what follows comes from that ranking.
The ladder
Five tiers of evidence
Each detector stamps its find with one of these. It is a closed list: nothing invents a sixth rung, and nothing promotes itself.
- 5 · Private keyIt announces itselfEverything between the opening and closing banners. The only tier whose default action is block.-----BEGIN …
- 4 · ProviderA company stamps its keysA literal prefix, a fixed alphabet, a known length.ghp_ · AKIA · sk-
- 3 · StructuralIt has a shape of its ownA grammar no vendor owns: headers, URLs, JWTs.Bearer …
- 2 · ContextualSomething called it a secretA credential-ish name, then a random-ish value.api_key = …
- 1 · EntropyJust looks randomNever enough on its own. A supporting signal only.x8Kd92mQz1
↑ More certainWeakest ↓
Tier 4
A prefix, an alphabet, a length
53 providers, 108 credential families stamp their keys with a recognisable opening. Every one of those detectors has the same four-part shape, and the whole tier is written in this small language.
- Prefix
- ghp_
- one literal string
- Alphabet
- A–Z a–z 0–9
- one of 11 fixed byte classes
- Run
- exactly 36
- exact n, or at-least n
- Validator
- none
- optional extra check
- Boundary
- The character just before and just after must not be part of the alphabet, so a slice of a longer blob is not mistaken for a whole key.
- Case
- A vendor documented as lowercase hex uses a lowercase-only class, so a case-mangled lookalike is rejected rather than matched.
- Validator
- Confluent's keys carry a real CRC-32 checksum tail, which the detector recomputes. Cloudflare's is only checked for being lowercase hex.
Two vendors are quirkier than the shape allows and get hand-composed grammars:
sk-proj-⟨74 chars⟩T3BlbkFJ⟨74 chars⟩That marker in the middle is base64("OpenAI"). Every key they issue carries it, so the detector anchors on it instead of trusting sk- alone.
Tier 3
Shapes nobody owns
These are general formats that no vendor owns. There are 6 of them, and they are the only detectors the small common profile keeps.
- JWT
- Three dot-separated chunks, and the first two must start
eyJ, the fingerprint of base64'd JSON. - Bearer
Bearerplus at least 16 characters. Only the token is selected, never the header word.- Connection URL
- Picks out only the password inside
postgres://user:pw@host. The host stays readable. - otpauth URI
- The scheme plus a
secret=parameter. The label and issuer are left alone.
There is one deliberate exception: a JWT is dropped if its payload says it is a Supabase anon key, which is public by design and ships in browser bundles. It is the only place a token's contents are read at all.
Tier 2
The name vouches for the value
This is the catch-all for credentials nobody has a grammar for. It needs two things to agree: a name that sounds like a secret, and a value that looks like one.
- Name
- Case and punctuation are flattened first, so
apiKey,API-KEYandapi.keyare one name. - Strong names
api_key,password,client_secret,access_token: the value must score 3.0 on randomness.- Maybe names
auth,credential,signing_key: the bar rises to 3.5. Weaker word, stronger proof.- Length
- At least 8 characters; randomness is only measured past 16; anything over 4 KB is left to the specialists.
- Plain token
- The bare word
tokenis deliberately ignored. Too common to mean anything.
Then it throws most of them away
A large part of this detector is knowing what is not a secret. All of these are recognised and skipped:
- changeme
- placeholder
- redacted
- <your-key-here>
- ${process.env.KEY}
- $[variables.x]
- `date +%s`
- op://vault/item/field
- :bind_param
- /etc/ssl/key.pem
- true
- 12345
Every one of those is a pointer to a secret, not the secret. Flagging them is the fastest way to make a scanner people turn off.
Tier 1
Randomness is not evidence
Not flagged
a8f3k29dj4ms91x alone in your text. A commit hash, an ID, a nonce, a UUID.
Flagged
password=a8f3k29dj4ms91x. The same string, now with something standing behind it.
Entropy never promotes anything by itself.
It only raises or lowers the bar for a tier above it. That single decision is why this library reports so few false alarms, and why a high-entropy secret with no name and no shape walks straight past.
When tiers disagree
Five tie-breakers, always in this order
Two detectors often claim overlapping text. The winner does not depend on which ran first. A fixed comparison runs top to bottom until something differs.
- 1 · Severity
- A candidate that would only warn can never displace one that would block. This outranks even the tier.
- 2 · Tier
- The ladder above. Private key beats provider beats structural beats contextual beats entropy.
- 3 · Confidence
- High, medium, low.
- 4 · Width
- The narrower span wins. Redact the key, not the paragraph around it.
- 5 · Order
- Registration order, then emission order. Built-ins register before anything custom, so a stable answer always exists.
Across the whole input it does not take winners greedily: it picks the combination of non-overlapping finds with the highest total evidence, weighted by these same keys.
The little language
There is no regex engine
The core is allowed zero outside dependencies, so it has no regex engine to use. Instead the four-part shape from tier 4 was written down once, and every provider detector is an instance of it.
What that buys
No backtracking exists, so the classic one-evil-input-hangs-the-server attack has nowhere to live, by construction rather than by review. A lookup table computed once per scan keeps the whole pass linear, even on hostile input.
What it costs
Anything that is not “prefix, then a bounded run” needs hand-written code. Two vendors' formats needed exactly that.
The 11 alphabets, and nothing else
- alnum
- alnum-dash
- alnum-dash-dot
- alnum-underscore
- upper-alnum
- lower-alnum
- digit
- hex
- hex-or-dash
- lower-hex
- base64-body
Bring your own rules
Your format, their engine
Shipped · Rust · JavaScript · Python · CLI
Companies have their own key formats. Adding a detector used to mean writing Rust, so everyone else ran a second scan of their own beside this one, which is exactly the divergence the library exists to prevent. Now you hand over the same four-part shape as plain text, as data rather than a callback.
ruleset-revision: 1
detector: acme-internal-token
specificity: contextual
prefix: "ACME_"
alphabet: alnum-dash
run: at-least 20
validator: noneA callback, still rejected
Your code runs on every candidate. It needs the plaintext to decide, so plaintext crosses the boundary, and Node, the browser and Python could each answer differently.
A ruleset, what shipped
Your grammar crosses the boundary once, up front. The Rust core does every match itself, through the engine it already has. Nothing to execute, so nothing can capture your text.
- Tier cap
- A rule may claim entropy or contextual only. The top three tiers are reserved for built-ins, so your rule can add finds, never overturn one.
- Validators
- Chosen by name from a closed list. Never a function, an expression, or a body you write.
- Confidence
- Always medium. There is no field to raise it.
- Size
- 64 KiB of UTF-8, parsed once up front. Fail closed: one bad line rejects the file.
Personal data
A sixth thing to find, kept apart
Personal data has its own domain beside the five tiers: email, IBAN, payment card, phone, network address and US social security number. It is opt in and off by default on every surface, and activation is explicit.
Luhn
Payment cards are validated by the checksum the card industry already uses, as a named, versioned validator.
IBAN mod-97
Bank account numbers are checked by the ISO remainder test, not by shape alone.
SSN allocation
Structural exclusions the US agency publishes, never issuance or identity lookup.
These three are the first arithmetic checks in the engine. Before them, the only numeric test in the whole core was one CRC-32 on one provider's key.