Skip to content

AI Agents & Cybersecurity Editorial Strategy


The mission of Hermes Codex is to establish a rigorous, evidence-grounded reference on the empirical capabilities and physical limits of AI agents interacting with real-world software, networks, and binary environments.

The central analytical framework follows an unbending sequence:

Capability ───> Benchmark ───> Experiment ───> Limitations ───> Operational Implications
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ HERMES CODEX β”‚
β”‚ Empirical AI Agent Cybersecurity Knowledge Architecture β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β”‚
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β–Ό β–Ό β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ CLUSTER A β”‚ β”‚ CLUSTER B β”‚ β”‚ CLUSTER C β”‚
β”‚ AI AGENTS THAT HACK β”‚ β”‚AI AGENTS THAT GET HACKEDβ”‚ β”‚ AI AGENTS THAT DEFEND β”‚
β”‚ (Offensive Frontier) β”‚ β”‚ (Systemic Exposure) β”‚ β”‚ (Autonomous Defense) β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€ β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€ β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ β€’ Vulnerability Hunting β”‚ β”‚ β€’ Prompt Injections β”‚ β”‚ β€’ SOC Triage & Alerting β”‚
β”‚ β€’ PoC Generation β”‚ β”‚ β€’ Tool & MCP Poisoning β”‚ β”‚ β€’ Malware Analysis β”‚
β”‚ β€’ Exploit Synthesis β”‚ β”‚ β€’ Memory / RAG Tamper β”‚ β”‚ β€’ Automated Patching β”‚
β”‚ β€’ Binary Reverse Eng. β”‚ β”‚ β€’ Lateral Privilege Esc β”‚ β”‚ β€’ Threat Intel CTI β”‚
β”‚ β€’ Autonomous Pentesting β”‚ β”‚ β€’ Agent-to-Agent Spoof β”‚ β”‚ β€’ Cyber-Range Sim β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Rather than treating β€œAI hacking” as a binary capability, Hermes evaluates autonomous agents along an empirical 10-level capability ladder:

LevelCapability MilestoneVerification ThresholdPrimary Benchmark
L1Vulnerability ComprehensionAccurately identifies root-cause CWE and affected lines in isolated snippetsHumanEval-Sec, SecFix
L2Vulnerability DiscoveryIdentifies unpatched flaws across complex multi-file codebases without hintsCyberGym, SecCode
L3Bug ReproductionTakes a vulnerability report/CVE and reproduces the crash in a containerCyberGym, SRE-Bench
L4PoC GenerationGenerates an input triggering an unhandled exception or denial-of-serviceExploitGym (PoV phase)
L5Exploitation Primitive SynthesisObtains controlled arbitrary read, arbitrary write, or out-of-bounds corruptionExploitBench
L6Control-Flow HijackingBypasses local protections (stack cookies, basic canary) to redirect executionExploitBench, ExploitGym
L7Arbitrary Code Execution (ACE)Gains unauthenticated arbitrary shellcode execution (with ASLR / DEP disabled)ExploitGym (Userspace)
L8Mitigation Evasion & Full RCEBypasses modern mitigations (ASLR, DEP, CFI) to obtain remote executionExploitGym (Kernel / V8)
L9Session Persistence & Post-ExDiscovers internal credentials, corrupts execution environments, maintains accessInter-Agent Benchmarks
L10Autonomous Multi-Stage CampaignExecutes recon $\to$ weaponization $\to$ lateral movement across an enterprise networkCyber-Range Simulation

Every research analysis published within this series adheres to an invariant, peer-review grade structure:

  1. Introduction: Immediately presents the core empirical question and real-world stakes.
  2. Why This Paper Matters: Formulates the architectural problem and prior state of knowledge.
  3. How the Research Works: Dissects the experimental methodology, tooling, and execution testbeds.
  4. Dataset & Benchmark Architecture: Quantifies task counts, line-of-code scales, anti-analysis primitives, and contamination defenses.
  5. Models Tested: Identifies model versions, inference parameters, tool APIs, and reasoning scaffolds.
  6. Empirical Results: Tabulates exact pass rates, execution costs, single-pass vs. multi-turn loops.
  7. Critical Analysis: Unpacks methodological biases, evaluation leaks, and real-world divergences.
  8. What Can an AI Agent Actually Do? (Mandatory): Explicitly partitions findings into Demonstrated Capability, Reasoned Inference, and Hypothetical Speculation.
  9. Offensive Implications: Analyzes actionable impact for red teams, penetration testers, and threat actors.
  10. Defensive Implications: Identifies actionable controls for detection engineers, SOCs, and security architects.
  11. Next 12–24 Months Outlook: Grounded, non-sensationalist predictive vectors.
  12. Related Intelligence & Links: Direct bidirectional graph connections to CVEs and Agentic Attack Patterns.

Incoming preprints across cs.CR, cs.AI, cs.SE, and cs.LG are evaluated against a standardized 100-point rubric:

Scientific Value (25 Pts)

Methodological rigor, contamination avoidance, statistical significance, and artifact reproducibility (open code, datasets, containers).

Cybersecurity Relevance (25 Pts)

Direct alignment with real-world operating systems, kernel runtimes, active protocols, and physical threat surfaces.

Traffic & Search Intent (25 Pts)

High organic search volume, clear practitioner queries (β€œCan AI reverse engineer binaries?”), and enduring reference utility.

Editorial Potential (15 Pts)

Clarity of thesis, potential to illuminate complex architectural boundaries, and modularity within our content clusters.

Longevity (10 Pts)

Long-term relevance as a baseline benchmark or paradigm shift; resistant to instant model-version obsolescence.

Threshold: Papers scoring >= 75 / 100 qualify for full deep-dive publication.


The initial wave of 11 publications establishes Hermes Codex as the benchmark-of-record:

  1. Article 1: Can AI Agents Really Perform Reverse Engineering? β€” SRE-Bench
  2. Article 2: Can an AI Agent Really Exploit a Real Vulnerability? β€” ExploitGym
  3. Article 3: AI Cybersecurity Benchmarks Compared: CyberGym vs ExploitGym vs ExploitBench vs SRE-Bench
  4. Article 4: ExploitBench: How Far Can an AI Agent Go When Exploiting a Vulnerability?
  5. Article 5: Can One AI Hack Another AI? β€” PIMiner & Agent Against Agent
  6. Article 6: Are Prompt Injections Impossible to Solve? β€” Fundamental Limits of Context Separation
  7. Article 7: Is MCP a Major New Attack Surface for AI Agents? β€” Hybrid Analysis & MTGuard
  8. Article 8: What Can AI Agents Actually Do in Cybersecurity in 2026? (The Master Pillar Page)
  9. Article 9: Can LLMs Accelerate Cyber Ranges and Attack Simulations? β€” arXiv:2608.16422
  10. Article 10: The State of Generative AI in Cybersecurity & Privacy: 2026 Landscape, Frontiers, and Gaps β€” arXiv:2607.06963
  11. Article 11: Can You Trust LLMs with Cyber Threat Intelligence? Empirical Limits and Hallucinations β€” arXiv:2503.23175

Hermes maintains a living empirical index tracking demonstrated agent capabilities across release cycles:

Capability Domain2024 (GPT-4 / Claude 3)2025 (o1 / Sonnet 3.5)2026 (Frontier Reasoning)Verified Frontier Baseline
Vulnerability Discovery (Source)MediumHighHighSnyk / Semgrep AI
PoC Crash GenerationLowMediumHighCyberGym / ExploitGym
Binary Reverse EngineeringVery LowLowLow / MediumSRE-Bench (arXiv:2608.11469)
Kernel / V8 ExploitationNegligibleVery LowLowExploitGym (arXiv:2605.11086)
Automated Prompt InjectionLowMediumHighPIMiner (arXiv:2608.05108)
Agent Tool SandboxingLowLowMediumMTGuard (arXiv:2607.25297)
CTI Report AnalysisLow / HallucinatoryLowMediumarXiv:2503.23175