Blog
Research, benchmarks, and field notes.
How RedPick performs against public benchmarks, what we find in real applications, and where AI-driven offense and defense are heading.
HackTheAgent 5/5: RedPick on the HackAIcon AI-Agent CTF
RedPick solved all 5 HackTheAgent challenges server-confirmed — prompt injection, refusal mining, agentic business-logic abuse, and an SSRF redirect to an internal endpoint. A live LLM-agent CTF, no source code.
READ THE ARTICLERedPick on Pensar's Argus Benchmark: 57/60 Black-Box + 3 Runtime Fixes
RedPick captured 57/60 flags on Pensar's Argus black-box benchmark, with three runtime-blocked labs validated after narrow upstream patches (60/60 patched) — raw and patched results reported separately, on purpose.
READ THE ARTICLE20/20 on Duck Store: RedPick on Escape's Agentic Pentesting Benchmark
RedPick found all 20 known vulnerabilities on Escape's Duck Store benchmark — source-free grey-box — plus 88 validated extras beyond the answer key, at 95.6% extended precision.
READ THE ARTICLE55 Findings, No Source Code: RedPick on Doyensec's Aikido vs XBOW Apps
RedPick ran Doyensec's two Aikido-vs-XBOW apps, Photoview and Fider, fully black-box: 55 extended true positives and 34 valid findings beyond the answer key — no source code.
READ THE ARTICLE16/16 on HackBench: From UNION SQLi to 3-Stage RCE
RedPick achieved a perfect 16/16 (4000/4000 pts) on HackBench, a collection of 16 real-world CVE-based web exploitation challenges spanning SQL injection, XSS, authentication bypass, IDOR, command injection, and multi-stage RCE chains. Fully automated, no human intervention.
READ THE ARTICLEProjectDiscovery Vibe-Coding: 152 Findings, 0 FP
RedPick reached 74/74 ground-truth coverage on ProjectDiscovery's Vibe-Coding Benchmark, plus 78 additional code-backed vulnerabilities — 152 total at 100% precision across vaultbank, medportal, and claimflow, with 10 end-to-end attack chains and zero false positives.
READ THE ARTICLEThe AI Attacker Era: Why Defenders Need AI Too
AI agents are now finding zero-days, exploiting one-day vulnerabilities, and reshaping how offensive security works. Defenders relying on yesterday's tools are testing against yesterday's attackers.
READ THE ARTICLERedPick scores 7/7 on HackMerlin — A 4-Layer LLM Defense Cracked
RedPick achieved 7/7 on HackMerlin, a progressive LLM prompt injection challenge. Walkthrough of the Cloze Filter Detection technique that bypassed a 4-layer defense fully automated, no human intervention.
READ THE ARTICLE100% on PortSwigger Academy — 274 Labs + 5 Mystery
RedPick solved all 274 PortSwigger labs and 5 mystery labs — 31 vulnerability categories including the new AI-powered scanner labs, fully automated, zero human intervention.
READ THE ARTICLEAgentic vs Automated: Why Traditional DAST Falls Short
There's a fundamental difference between automation and agency. This is what it means for application security, which vulnerabilities DAST will never find, and when traditional scanners still make sense.
READ THE ARTICLELLM Security Testing: The New Attack Surface
Your AI features are an attack surface. Prompt injection, jailbreaking, data exfiltration — here's the full OWASP LLM Top 10, what you need to test, and how RedPick approaches LLM security.
READ THE ARTICLE