Proof

Five benchmarks. Five perfect scores.

RedPick was submitted to leading public benchmarks for autonomous penetration testing and scored 100% on all of them. No cherry-picking, no partial results, no asterisks. Every score below was achieved black-box, or source-free grey-box.

Benchmark methodology

Run under the strictest available conditions.

Black-box mode

No source code, no internal documentation, no hints provided to RedPick.

Fully autonomous

No human intervention during test execution. RedPick planned, executed, and verified each exploit independently.

Reproducible

These benchmarks are publicly available. Results can be independently verified by running the same suites.

No cherry-picking

Every challenge in every benchmark was attempted. Scores reflect the complete suite, not a selected subset.

FAQ

Common questions.

Can AI actually find real vulnerabilities?

Yes, and the results are public. RedPick scored 104/104 on XBOW, 274/274 on the PortSwigger Web Security Academy, 16/16 on HackBench, and 74/74 on ProjectDiscovery, all in black-box mode. Those cover the vulnerability classes that appear in real engagements: injection, broken access control, business logic, SSRF, and multi-stage exploitation chains.

Are the benchmark results independently verifiable?

Yes. These are public, third-party benchmarks — anyone can reproduce them. RedPick also publishes per-challenge walkthroughs and Tier-1 evidence on GitHub, so the findings can be reviewed and re-run, not just taken on trust.

How does RedPick compare to other autonomous pentest agents?

On XBOW's own benchmark RedPick is the only system to reach 104/104 (100%) in pure black-box, no-hint mode — ahead of the next published results, and without the source-code access that some higher-scoring white-box agents rely on. The full per-agent comparison is in the XBOW write-up.

Let's prove it on your application.

Benchmarks — Five Perfect Scores | RedPick