30/34
88.24%threshold 75%
Planted defects Caraxe proved it reached, with mechanical location or control evidence.
We hide real bugs in real apps and see how many Caraxe finds. Someone else holds the answers and does the scoring. Our most recent round is the first to clear every bar we set.
30/34
88.24%threshold 75%
Planted defects Caraxe proved it reached, with mechanical location or control evidence.
20/30
66.67%threshold 60%
Exact catches among the defects it proved it reached.
20/34
58.82%
End to end, including the cost of paths it never reached.
0/3
0.00%threshold at most 5%
Exact planted-class claims made against clean control builds. Generic suggestions are counted separately.
Our most recent round: 37 builds, 34 with a bug planted and 3 clean, 0 runs that failed to finish. Every number on this page is read out of that run’s own results file and checked against its hash before the page is built. Nobody types these in.
Caraxe is better on some UI frameworks than others, and that gap is wider than the gap between test rounds. So we show them separately.
| UI framework | Planted | Reached | Caught | Detection when reached | Overall |
|---|---|---|---|---|---|
| Android Views | 8 | 7 | 4 | 57.14% | 50.00% |
| Jetpack Compose | 13 | 11 | 7 | 63.64% | 53.85% |
| React Native | 13 | 12 | 9 | 75.00% | 69.23% |
Our last five rounds. Round two went backwards and it stays on the page. Round four beat everything before it and still failed, on one false alarm.
| Test round | Overall catch | Outcome |
|---|---|---|
| Round 1 | 26.47% | - |
| Round 2 | 20.59% | went backwards |
| Round 3 | 38.24% | - |
| Round 4 | 55.88% | failed on one false alarm |
| Most recent | 58.82% | passed |
Scoring marked 1 of 3 clean builds as a false alarm and withheld the pass. Then we read the code. The control Caraxe flagged pops a message and does nothing else – no navigation, no field, no saved state. Caraxe was right and the clean build was not clean.
The claim was reclassified as a real defect in a nominally clean fixture. No scorer, matcher or engine rule was changed to reach that conclusion. Both documents are published: the original score that says NOT ISSUED is preserved unedited beside the certified one. This is the third time a control built to be clean has turned out to contain a real bug.
UNSEAL_REPORT.md and UNSEAL_REPORT.raw.md.You can review the same app again. We tested whether the second run should be steered by what the first one found, and it was not better, so we do not do that. Past that, there is less we can currently prove than we would like, and this is the honest state of it.
No second run is steered by the first. We tested that and it did not beat simply running again, so we do not do it and do not claim it. That is the one result here that survived review.
We are not publishing a number for how much more a second run covers. We had one, and an audit showed the measurement counted the same screen twice whenever its contents changed. The coverage figure is withdrawn until we can count it properly.
We are not yet claiming that a second run re-checks what the first one found. The machinery that records those answers works; the step that recognises the same issue across two runs currently recognises nothing, for the same reason the coverage number was withdrawn. It goes live when that is fixed and a two-review test passes, and not before.
Planted bugs show whether Caraxe finds what we hid. Real apps show whether it says anything worth acting on. A second reviewer checked every finding here, and the strongest were re-tested on a real phone.
We wrote these rules down before we had a score, and we have not loosened them since.
Lock the build, the apps, the device and the scoring rules before anyone sees a result. Change any of them and it is a new test.
Test across several apps and UI frameworks, including flows that need navigation, saved state and real data.
Score against a locked answer key, and include clean builds so a high score cannot hide a noisy one.
Run it again on real devices. Say when a run did not finish, and track whether it reached the screen at all, not just whether it reported something.
Prove it finishes, costs what we said, keeps its evidence and can be rolled back, on the setup customers actually get.
Publish the misses, the exclusions and the per-app results next to the headline. Runs that proved nothing stay on the page.
We publish what it misses too. In our most recent round Caraxe missed 10 bugs on screens it reached, and never reached 4 more. Clearing our bar means it cleared three at once, not that detection is solved. This is 34 planted bugs across three UI frameworks, not a survey of every Android app, and accessibility findings point at published guidelines rather than legal sign-off.
We are working with a small number of Android teams whose real release workflows can strengthen the multi-app validation set. Participation is scoped before any review.