Skip to content
CAI
Software that uses CAICheck a score

The noise standard

Publish your false-positive rate, and let it be checked.

CAI scores a codebase from evidence gathered by each implementation. Its scoring rules do not assess the tools that gather that evidence. This companion standard sets out a method for measuring the noise rate of any static-analysis tool, whether or not it is used with CAI.

A vendor can quote a noise rate for its own scanner. Under this method, a holdout test set is drawn before any tool runs, findings are assigned verdicts from a fixed list, and every submission must meet the same requirements. The standard checks the run, calculates the rate from the submitted counts, and publishes the result regardless of how the tool performs. It does not run the scanner or decide whether individual findings are valid.

What the standard does

It verifies and publishes. It does not measure for you.

You measure

You run your tool over the published holdout, at the revisions it pins, and you judge your own findings however you judge them. The standard never sees your engine and takes no position on it.

You submit counts and the findings behind them, never a rate. A participant that computed its own headline would be marking its own homework.

The standard checks, then computes

Nine checks run against the submission: holdout membership, pinned sha, run ordering, claim class, recency declaration, configuration declaration, finding count, re-judge and coverage. Any one of them can refuse the run, and a refusal names what failed.

Then it computes the rate, its confidence interval, the actionability rate, the per-100k absolutes, the smallest difference the sample could detect, and two averages: one pooling every finding, and one giving each repository equal weight. All of it publishes together with the census it was drawn from.

The vocabulary

Six verdicts, because a binary hides two different failures.

Agreement between two tools answering different questions is not a measurement. The standard owns the verdict set, and every participant implements the same one.

Should not have fired

Untrue at the cited code, or not worth a reader's time. Scores as noise.

True, and actionable

True, and it says enough to act on. Scores as valid.

True, but too thin to act on

Scores as valid, and carries its failure on the actionability axis instead. A correct finding nobody can act on is a true positive for the detector and a failure for the reader. Forcing a rater to call it one or the other is how two careful people end up disagreeing about the same finding.

Both positions wrong

Neither the finding nor its opposite is right. Scores as noise, and escalates.

Cannot tell

The evidence shown was not enough to decide. That is a process defect rather than a verdict. From a human it leaves the rate alone; from a machine it must escalate, because excluding it there would hand a pipeline a way to duck its hardest cases and still report a clean number.

Rubric ambiguous

The rubric has no determinate answer here. Also a process defect: it leaves the rate and files against the method rather than the finding.

The holdout

Drawn before anyone runs, from a pool anyone can read.

The pool is public, or the draw proves nothing

Re-deriving a holdout needs the seed and the pool it was drawn from. Publishing only the seed proves nothing, because the pool could have been chosen afterwards to suit it. Both ship inside one signed manifest: the rules, the draws with their seeds and timestamps, and every candidate that was eligible.

Verify it yourself with openssl, against the key the manifest names. If the shipped manifest does not verify, the holdout endpoints answer 503 and serve no draw at all. A holdout that quietly degraded to a pool with no signature on it would be worse than one that stops.

Four repositories are never trained on

The recency strata compare never trained on against trained on N cycles ago. If every repository is eventually developed against, that first bucket empties and the comparison quietly stops meaning anything, so a reserved slice is drawn every period and kept untouched.

Every draw publishes its discards as well: drew 22, discarded 4, measured 18. Failing to build correlates with size and complexity, which plausibly correlates with noise, so a discard nobody reports is a quiet thumb on the scale.

What a submission must carry

Every field exists because leaving it out flattered somebody.

The findings, with coordinates

Repository, pinned sha, file and line wherever the dimension can supply them, rule id, title, and the claim class: whether the finding points at a specific location, at the shape of a named artefact, at risk, or at a recommendation resting on another finding. A claim about risk cannot be shown false at a line of code, so it carries no noise rate. A finding with no file, or a file with no line, cannot be matched against what another vendor reported, so the union counts it as unmatchable and says so rather than dropping it quietly.

The configuration

Which ruleset ran, whether that is the product default, and any rule disabled or threshold moved. Every other check constrains the run itself and none of them constrains how the tool was set up beforehand, and a result produced under a carefully tuned profile is not the result a customer gets out of the box.

The denominator

Production lines of code, because the absolute figures are what expose suppression: noise per 100k lines and valid findings per 100k lines. A tool that goes quiet by firing less often improves its rate and worsens its absolute, and only publishing both makes that visible.

Readable git history

A contained scan without a usable git directory makes the history-derived dimensions emit false verdicts. Their noise is then an environment artefact rather than a capability gap, so a run that cannot say whether history was readable is refused.

A recall counterpart

A noise rate measures precision and nothing else, and precision on its own rewards under-firing. Published alone it is an incentive to say less, so a rate arrives with a recall figure beside it, or with the reason there is none.

A census that balances

Reported equals adjudicated plus excluded plus unrated. The findings that fall out of a funnel are exactly the ones a reader wants to know about, so the exclusion rate is capped at 5 per cent. Above that the run is void rather than published with a caveat attached.

The mark

CAI-measured: free, mechanical, and never a certificate.

The mark says a measurement happened properly. It says nothing about how good the tool is: three of its four conditions are about process, so a noisy tool that ran the published draw, submitted in time and published in full earns it. The rate is what says how it went.

Every condition is a fact the standard already holds. A judgement call anywhere in here would be an unreviewable veto held by one participant over its rivals, and a withheld mark names the condition that failed.

ran-against-the-published-holdout

Decided by the membership and sha checks the receipt already reports.

submitted-before-the-deadline

A submission received before the period's published deadline.

the-run-reproduces

The receipt is accepted, which means every verification check passed.

numbers-published-in-full

A publication exists for the period and met the contract.

It is free, and that commitment is a rule about timing rather than an intention. Any change, whether a fee or a new condition, takes effect no earlier than one full published period after it is announced, and never for a period already opened. A change reaching an open period would bind participants who have already spent the compute and cannot withdraw.

The crowd

The one check that comes from outside the model family.

Four judges from one vendor agreeing tells you they are consistent. It does not tell you they are right, and adding more judges from the same family never converts the one into the other. So a sample of what the machines agreed on goes to people, not only the contested tail.

One item at a time

The finding, a link to the code at the revision that was scanned, and the six choices. Nothing about which tool reported it, and nothing about what the judges said. Tell somebody four judges already agreed and a reasonable person rubber-stamps, and the spot-check exists precisely to catch the case where all four were wrong together.

Nobody rates their own estate

"This isn't a real problem" is a very human reaction to your own code being criticised. It biases a rate systematically rather than randomly, and no amount of averaging removes a bias that leans the same way every time.

Calibration is earned, not agreed

Raters are scored against findings settled outside the rating process: an upstream fix that got merged, a vendor withdrawal, an advisory that was retracted. Scoring them against what the crowd agreed would measure conformity and then call it accuracy.

Scores publish, and are never applied

No answer is dropped for a poor calibration score. Excluding raters by the very thing being measured is selection on the outcome, and what it leaves behind is the subset that happened to agree: a cleaner number that means considerably less.

Taking part

Seven steps, and every one of them is an endpoint.

No account, no key, no fee. The API is open and rate-limited per client; the crowd's endpoints carry their own budget of one item a minute.

1 · Register intent

POST /api/noise/intent, before the period's holdout is drawn. It closes at the draw, because intent declared with the sample already in hand is not intent. The register publishes who said they would take part, so not submitting is visible too.

2 · Read the method

GET /api/noise/method returns the whole contract as JSON: the verdicts, the checks, the ceilings, the claim classes, the panel rule and the mark's conditions. The panel rule is that no model may appear twice in a panel, and that a panel must span at least two model families. It is machine-readable so your client can be built against the contract itself rather than against a PDF of it.

3 · Fetch the holdout

GET /api/noise/holdout/{period} returns the drawn repositories with their pinned revisions, the seed, and the signature over the corpus they came from. Verify that signature before you trust the draw.

4 · Run, at the pins

Your tool, your judging, your machine. The run starts after the draw was published, at the exact revisions it pins, with git history readable. A run that started before its own holdout existed is answering a different question.

5 · File the run

POST /api/noise/submissions carries the findings with their coordinates, the recency of each repository, and the configuration you ran under. The receipt names every check that passed or failed, and the findings it recorded.

6 · Publish the result

POST /api/noise/publication carries the counts, the claim-class breakdown, the per-repository tallies and the denominator. The standard computes the figures, and refuses the publication if the contract has not been met, naming each breach.

7 · Read the record

GET /api/noise/record/{period} and /api/noise/mark/{period} return the verdicts, the disputes raised against them, and who earned the mark. Your own register row stays hidden from everyone else until the period publishes.

The standard's maintainer also competes in it.

The Code Assurance Initiative publishes this standard. Its only member today is Canine Development, which also builds a scanner that takes part in this programme — so the conflict is exactly what it was before the Initiative existed, and moving the standard under a new name does not resolve it. It is not resolved by anybody promising to behave either. Four things make it survivable, and all four are mechanical:

  • Depth is measured leave-one-out. Each tool is scored against the union of what every other participant found, itself excluded. A designated baseline would turn depth into a measure of how closely you agree with whoever holds that position, and would score them 100 % by construction.
  • An embargo covers the register. No participant sees another's result before the period publishes, and the maintainer has no exemption. Early sight of a rival's number is the single most valuable thing this position could be worth.
  • The mark's conditions are facts, not judgements. Nothing in granting or withholding it is anyone's opinion, and a withheld mark names the condition that failed.
  • The method version takes effect from the next holdout drawn, never the current one. A change published after a draw cannot apply to that draw, including one the maintainer would benefit from.

This is phase one: a published method with an open invitation, and no governance body yet. The same question applies to the standard as a whole, and it is answered in the same place.

Run it against your own tool, and publish what you get.

A verdict you disagree with can be contested, with a reason, at /api/noise/verdicts/{findingId}/dispute. Upheld and overturned disputes both publish, because a contest that only surfaced when the challenger won would be a complaints box. The scoring standard itself → the standard.