Skip to content

AI Cybersecurity Benchmarks Compared: CyberGym vs ExploitGym vs ExploitBench vs SRE-Bench

Benchmark Suite4 Major Evaluation Frameworks
Total Evaluated Tasks4,000+ Deterministic Challenges
ScopeCrashes Β· Exploitation Β· Binaries
State of the ArtFrontier Reasoning Models (2026)

1. Introduction: The Fragmentation of AI Cybersecurity Evaluation

Section titled β€œ1. Introduction: The Fragmentation of AI Cybersecurity Evaluation”

Between 2024 and 2026, the cybersecurity evaluation of Large Language Models transitioned from static multiple-choice questionnaires (such as SecQA) to interactive, dynamic execution testbeds. However, because each research lab developed its own isolated benchmark, published claims frequently contradict one another:

  • One paper reports that AI models achieve an 82% success rate in cybersecurity tasks.
  • Another paper evaluating the same frontier model reports a failure rate exceeding 85%.

These conflicting headlines arise because β€œhacking” is not a monolithic activity. There is a vast technical gulf between generating a crash-inducing payload for an open-source library and synthesizing a multi-stage Return-Oriented Programming (ROP) exploit against a hardened browser engine.

This study presents an architectural comparison of the four definitive benchmarks defining AI agent cybersecurity capabilities in 2026: CyberGym, ExploitGym, ExploitBench, and SRE-Bench.


β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ THE AI CYBERSECURITY EVALUATION LANDSCAPE β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β”‚
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β–Ό β–Ό β–Ό β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ CyberGym β”‚ β”‚ ExploitGym β”‚ β”‚ ExploitBench β”‚ β”‚ SRE-Bench β”‚
β”‚(arXiv:2506.02548β”‚ β”‚(arXiv:2605.11086β”‚ β”‚(arXiv:2605.14153β”‚ β”‚(arXiv:2608.11469β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€ β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€ β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€ β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ β€’ Focus: Crash β”‚ β”‚ β€’ Focus: Full β”‚ β”‚ β€’ Focus: 16-tierβ”‚ β”‚ β€’ Focus: Binary β”‚
β”‚ Reproduction β”‚ β”‚ Exploitation β”‚ β”‚ Capability β”‚ β”‚ Reverse Eng. β”‚
β”‚ & PoC Inputs β”‚ β”‚ & RCE Shells β”‚ β”‚ Ladder (V8) β”‚ β”‚ & Anti-Debug β”‚
β”‚ β€’ 1,507 Flaws β”‚ β”‚ β€’ 898 Flaws β”‚ β”‚ β€’ 41 Hardened β”‚ β”‚ β€’ 19 Programs β”‚
β”‚ β€’ 188 Repos β”‚ β”‚ β€’ Userspace, β”‚ β”‚ V8 Instances β”‚ β”‚ β€’ 1,572 Tasks β”‚
β”‚ β€’ Source Code β”‚ β”‚ V8 & Kernel β”‚ β”‚ β€’ Granular β”‚ β”‚ β€’ Stripped β”‚
β”‚ Available β”‚ β”‚ β€’ Live Flag Or. β”‚ β”‚ Primit. Flags β”‚ β”‚ Zero Contam. β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

The following table contrasts the design methodologies, execution models, and scoring mechanisms of all four benchmarks:

Evaluation DimensionCyberGymExploitGymExploitBenchSRE-Bench
Primary ObjectivePoC generation & crash reproductionArbitrary code execution (ACE/RCE)Exploitation pipeline stage verificationStripped binary analysis & anti-debug defeat
Task Volume1,507 CVEs898 Instances41 Targets (656 Tasks)262 Binaries (1,572 Tasks)
Target Surface188 C/C++ open-source projectsUserspace apps, V8 engine, Linux kernelGoogle Chrome V8 JIT engineNetwork daemons, firmware, malware, parsers
Input ArtifactsSource code, patches, issue reportsSource code, PoV crash input, compiler flagsV8 d8 binary, regression test case, diffStripped compiled ELF/PE binaries ONLY
Success VerificationASan crash trace / exit code matchDeterministic flag capture via shellcode16 automated challenge-response oraclesExact crypto key, patch, or protocol token
Contamination RiskHigh (Public GitHub CVEs & issues)Moderate (Historical CVEs with modified envs)Moderate (Known bugs, randomized oracles)Zero (100% private, newly authored code)
Frontier Model Pass Rate48.2% – 62.4%13.4% – 17.5%21.8% (Tier avg.)11.2% – 14.6%

CyberGym established the first large-scale pipeline for measuring whether AI models could take a known vulnerability advisory and generate a functional crash input.

  • Strengths: Huge sample size (1,507 vulnerabilities); tests diverse software architectures (image libraries, network tools, database servers).
  • Weaknesses: Uses crash reproduction as a proxy for exploitation. A null-pointer dereference that terminates a process is graded as a success, even though it provides zero exploitation value to an attacker beyond a simple denial-of-service.

ExploitGym directly tackles the weaponization bottleneck: given a crash, can an agent write an exploit that executes shellcode?

  • Strengths: Evaluates modern defensive mitigations (ASLR on vs. off); measures three distinct target domains including the Linux kernel.
  • Weaknesses: High resource cost; binary pass/fail grading hides partial progress along the exploitation path.

Rather than treating exploitation as an all-or-nothing milestone, ExploitBench decomposes the process into 16 granular stages:

  • The 16 Primitives: Crash reproduction $\to$ memory leak $\to$ fake object synthesis $\to$ arbitrary read $\to$ arbitrary write $\to$ control-flow hijack $\to$ sandbox escape.
  • Value: Reveals exactly where models stall. In V8 testing, models frequently achieve primitive 5 (arbitrary read) but fail at primitive 12 (sandbox isolation bypass).

SRE-Bench evaluates the hardest challenge in cybersecurity: understanding closed-source, compiled software stripped of all metadata.

  • Key Innovation: Zero dataset contamination. Every target program was authored from scratch in private repositories.
  • The Takeaway: Shows that when source code is removed and anti-analysis protections are applied, AI model effectiveness drops below 15%.

5. What Can an AI Agent Actually Do? (Cross-Benchmark Synthesis)

Section titled β€œ5. What Can an AI Agent Actually Do? (Cross-Benchmark Synthesis)”

By synthesizing the results across all four benchmarks, Hermes establishes a consolidated capability boundary:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ CONSOLIDATED CROSS-BENCHMARK CAPABILITY MAP (2026) β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ HIGH CAPABILITY (>50% Success Rate): β”‚
β”‚ βœ” Identify known vulnerability classes in open-source C/C++ (CyberGym). β”‚
β”‚ βœ” Generate fuzzing inputs that trigger existing ASan crashes (CyberGym). β”‚
β”‚ βœ” Decompile and summarize un-obfuscated x86-64 functions (SRE-Bench). β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ MODERATE CAPABILITY (15% – 35% Success Rate): β”‚
β”‚ β—‘ Exploit basic userspace stack overflows with ASLR disabled (ExploitGym). β”‚
β”‚ β—‘ Synthesize memory leak primitives to recover canaries (ExploitBench). β”‚
β”‚ β—‘ Solve simple crackmes under 5K LOC without anti-debugging (SRE-Bench). β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ LOW / NEGLIGIBLE CAPABILITY (<5% Success Rate): β”‚
β”‚ βœ– Autonomous weaponization of Linux kernel vulnerabilities (ExploitGym). β”‚
β”‚ βœ– Full multi-stage sandbox escapes in hardened JIT engines (ExploitBench). β”‚
β”‚ βœ– End-to-end reverse engineering of commercial-grade obfuscated binaries. β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

When an AI vendor claims their agent possesses β€œcybersecurity capabilities,” security leaders should ask three qualifying questions:

  1. Was the model tested on Source Code or Compiled Binaries?
    If the benchmark provided C source code (like CyberGym), the agent was doing code reasoning, not binary analysis.
  2. What was the Verification Oracle?
    Did the agent merely produce a segmentation fault (PoC crash), or did it execute arbitrary payload logic (ExploitGym)?
  3. How was Contamination Controlled?
    Were the test cases public CTF challenges existing on GitHub prior to the model’s training cutoff date?