Blog

Research, benchmarks, and field notes.

How RedPick performs against public benchmarks, what we find in real applications, and where AI-driven offense and defense are heading.

BENCHMARK

HackTheAgent 5/5: RedPick on the HackAIcon AI-Agent CTF

RedPick solved all 5 HackTheAgent challenges server-confirmed — prompt injection, refusal mining, agentic business-logic abuse, and an SSRF redirect to an internal endpoint. A live LLM-agent CTF, no source code.

June 10, 20268 min read
READ THE ARTICLE
BENCHMARK

RedPick on Pensar's Argus Benchmark: 57/60 Black-Box + 3 Runtime Fixes

RedPick captured 57/60 flags on Pensar's Argus black-box benchmark, with three runtime-blocked labs validated after narrow upstream patches (60/60 patched) — raw and patched results reported separately, on purpose.

June 8, 20267 min read
READ THE ARTICLE
BENCHMARK

20/20 on Duck Store: RedPick on Escape's Agentic Pentesting Benchmark

RedPick found all 20 known vulnerabilities on Escape's Duck Store benchmark — source-free grey-box — plus 88 validated extras beyond the answer key, at 95.6% extended precision.

June 6, 20268 min read
READ THE ARTICLE
BENCHMARK

55 Findings, No Source Code: RedPick on Doyensec's Aikido vs XBOW Apps

RedPick ran Doyensec's two Aikido-vs-XBOW apps, Photoview and Fider, fully black-box: 55 extended true positives and 34 valid findings beyond the answer key — no source code.

June 3, 202614 min read
READ THE ARTICLE
BENCHMARK

16/16 on HackBench: From UNION SQLi to 3-Stage RCE

RedPick achieved a perfect 16/16 (4000/4000 pts) on HackBench, a collection of 16 real-world CVE-based web exploitation challenges spanning SQL injection, XSS, authentication bypass, IDOR, command injection, and multi-stage RCE chains. Fully automated, no human intervention.

April 18, 202613 min read
READ THE ARTICLE
BENCHMARK

ProjectDiscovery Vibe-Coding: 152 Findings, 0 FP

RedPick reached 74/74 ground-truth coverage on ProjectDiscovery's Vibe-Coding Benchmark, plus 78 additional code-backed vulnerabilities — 152 total at 100% precision across vaultbank, medportal, and claimflow, with 10 end-to-end attack chains and zero false positives.

April 11, 202613 min read
READ THE ARTICLE
INSIGHTS

The AI Attacker Era: Why Defenders Need AI Too

AI agents are now finding zero-days, exploiting one-day vulnerabilities, and reshaping how offensive security works. Defenders relying on yesterday's tools are testing against yesterday's attackers.

April 10, 20269 min read
READ THE ARTICLE
BENCHMARK

RedPick scores 7/7 on HackMerlin — A 4-Layer LLM Defense Cracked

RedPick achieved 7/7 on HackMerlin, a progressive LLM prompt injection challenge. Walkthrough of the Cloze Filter Detection technique that bypassed a 4-layer defense fully automated, no human intervention.

April 9, 202611 min read
READ THE ARTICLE
BENCHMARK

100% on PortSwigger Academy — 274 Labs + 5 Mystery

RedPick solved all 274 PortSwigger labs and 5 mystery labs — 31 vulnerability categories including the new AI-powered scanner labs, fully automated, zero human intervention.

April 7, 202613 min read
READ THE ARTICLE
GUIDES

Agentic vs Automated: Why Traditional DAST Falls Short

There's a fundamental difference between automation and agency. This is what it means for application security, which vulnerabilities DAST will never find, and when traditional scanners still make sense.

March 10, 20269 min read
READ THE ARTICLE
RESEARCH

LLM Security Testing: The New Attack Surface

Your AI features are an attack surface. Prompt injection, jailbreaking, data exfiltration — here's the full OWASP LLM Top 10, what you need to test, and how RedPick approaches LLM security.

March 1, 202617 min read
READ THE ARTICLE

Ready to see what RedPick finds?

Blog | RedPick