Benchmark results

We published the standard first.
Here is the score.

We hide real bugs in real apps and see how many Caraxe finds. Someone else holds the answers and does the scoring. Our most recent round is the first to clear every bar we set.

Reach rate

30/34

88.24%threshold 75%

Planted defects Caraxe proved it reached, with mechanical location or control evidence.

Detection when reached

20/30

66.67%threshold 60%

Exact catches among the defects it proved it reached.

Overall catch rate

20/34

58.82%

End to end, including the cost of paths it never reached.

Clean false-credit rate

0/3

0.00%threshold at most 5%

Exact planted-class claims made against clean control builds. Generic suggestions are counted separately.

Our most recent round: 37 builds, 34 with a bug planted and 3 clean, 0 runs that failed to finish. Every number on this page is read out of that run’s own results file and checked against its hash before the page is built. Nobody types these in.

Where it is weakest

One number would hide this.

Caraxe is better on some UI frameworks than others, and that gap is wider than the gap between test rounds. So we show them separately.

UI frameworkPlantedReachedCaughtDetection when reachedOverall
Android Views87457.14%50.00%
Jetpack Compose1311763.64%53.85%
React Native1312975.00%69.23%
The record

Including the round that got worse.

Our last five rounds. Round two went backwards and it stays on the page. Round four beat everything before it and still failed, on one false alarm.

Test roundOverall catchOutcome
Round 126.47%-
Round 220.59%went backwards
Round 338.24%-
Round 455.88%failed on one false alarm
Most recent58.82%passed
Correction, published with the score

The first score said no.

Scoring marked 1 of 3 clean builds as a false alarm and withheld the pass. Then we read the code. The control Caraxe flagged pops a message and does nothing else – no navigation, no field, no saved state. Caraxe was right and the clean build was not clean.

The claim was reclassified as a real defect in a nominally clean fixture. No scorer, matcher or engine rule was changed to reach that conclusion. Both documents are published: the original score that says NOT ISSUED is preserved unedited beside the certified one. This is the third time a control built to be clean has turned out to contain a real bug.

A reader who sees only the certificate, or only the report that refuses it, has half the record. Both ship: UNSEAL_REPORT.md and UNSEAL_REPORT.raw.md.
Running it again

What we can and cannot say about a second run.

You can review the same app again. We tested whether the second run should be steered by what the first one found, and it was not better, so we do not do that. Past that, there is less we can currently prove than we would like, and this is the honest state of it.

No second run is steered by the first. We tested that and it did not beat simply running again, so we do not do it and do not claim it. That is the one result here that survived review.

We are not publishing a number for how much more a second run covers. We had one, and an audit showed the measurement counted the same screen twice whenever its contents changed. The coverage figure is withdrawn until we can count it properly.

We are not yet claiming that a second run re-checks what the first one found. The machinery that records those answers works; the step that recognises the same issue across two runs currently recognises nothing, for the same reason the coverage number was withdrawn. It goes live when that is fixed and a two-review test passes, and not before.

Beyond the planted set

45 findings on 22 apps nobody planted anything in.

Planted bugs show whether Caraxe finds what we hid. Real apps show whether it says anything worth acting on. A second reviewer checked every finding here, and the strongest were re-tested on a real phone.

That review removed our own headline. A crash we had reported as the most serious finding did not happen again in ten attempts on the same build, and no crash record was ever written, so it was withdrawn rather than published with a caveat. Every finding the gallery once called critical is gone the same way. 14 findings were withdrawn in total and 14 are held pending a device test.
The gate

What had to be true before we published this.

We wrote these rules down before we had a score, and we have not loosened them since.

Gate 01

Freeze what is being tested

Lock the build, the apps, the device and the scoring rules before anyone sees a result. Change any of them and it is a new test.

Gate 02

Use more than one kind of app

Test across several apps and UI frameworks, including flows that need navigation, saved state and real data.

Gate 03

Score catches and clean controls

Score against a locked answer key, and include clean builds so a high score cannot hide a noisy one.

Gate 04

Repeat the runs on devices

Run it again on real devices. Say when a run did not finish, and track whether it reached the screen at all, not just whether it reported something.

Gate 05

Prove the operating path

Prove it finishes, costs what we said, keeps its evidence and can be rolled back, on the setup customers actually get.

Gate 06

Publish the misses with the score

Publish the misses, the exclusions and the per-app results next to the headline. Runs that proved nothing stay on the page.

Limits of this result

What this does not say.

We publish what it misses too. In our most recent round Caraxe missed 10 bugs on screens it reached, and never reached 4 more. Clearing our bar means it cleared three at once, not that detection is solved. This is 34 planted bugs across three UI frameworks, not a survey of every Android app, and accessibility findings point at published guidelines rather than legal sign-off.

Publication record

What each result must carry.

  • The cohort, denominator and exclusions beside every summary number.
  • Per-family results, so an aggregate cannot hide a weak surface.
  • Every miss, kept in the record with the catches.
  • The artifact hashes and engine pin behind each figure.
  • Corrections published additively, with the superseded original preserved.

Help make the evidence representative.

We are working with a small number of Android teams whose real release workflows can strengthen the multi-app validation set. Participation is scoped before any review.