Skip to content

SysEvolve: Autonomous Attack-Defense Co-Evolution in Cyber Ranges β€” Analyzing arXiv:2608.15012

Paper ReferencearXiv:2608.15012
Research TeamsPeking Univ. / Southeast Univ. / APT Labs
Offensive EngineSYSSPEAR (>25% Pass Rate Boost)
Defensive EngineSYSARMOR (10–1000Γ— Precision)
Range Scale1,148 Ranges / 257 CVEs
Telemetry Overhead2.1% (Zero-Loss Tracing)

1. Executive Summary: The Cybersecurity Evolutionary Deadlock

Section titled β€œ1. Executive Summary: The Cybersecurity Evolutionary Deadlock”

For decades, cybersecurity has suffered from a structural asymmetry: attack automation accelerates exponentially, while defensive operations remain largely manual, reactive, and constrained by human analyst cognitive throughput. The emergence of frontier Large Language Models (LLMs) threatened to widen this divide. Attackers can leverage autonomous agents to rapidly craft polymorphic exploits, discover chainable misconfigurations, and navigate network perimeters.

However, existing evaluations of AI in offensive security have remained fundamentally disconnected from operational reality:

  1. Single-Target Myopia: Standard benchmarks like ExploitGym and ExploitBench evaluate agents in isolated single-container environments with predetermined vulnerability targets. They fail to test whether an agent can perform multi-host lateral pivots, identify intermediate credentials, or maintain state persistence.
  2. Lack of Dynamic Confrontation: Static synthetic evaluations pit an agent against a passive environment. Real-world adversaries encounter active Security Operations Centers (SOCs), Adaptive Endpoint Detection and Response (EDR) systems, and defensive honeytokens.
  3. Safety & Containment Hazards: Evaluating autonomous offensive agents on live systems introduces catastrophic runaway risks, including unsanctioned network exfiltration and self-propagating payload chains (AAP-007: Autonomous Cascading RCE).

In August 2026, a collaboration between Peking University, Southeast University, and enterprise threat intelligence labs introduced SysEvolve (arXiv:2608.15012): the first AI-native, autonomous, safe, adversarial attack-defense co-evolutionary system. Operating across 1,148 virtualized multi-host cyber ranges and encompassing 257 real-world CVEs, SysEvolve demonstrates how offensive AI (SYSSPEAR) and defensive AI (SYSARMOR) can continuously stress-test and evolve one another within a strictly contained declarative orchestration substrate (SYSFIELD).

THE SYSEVOLVE CLOSED-LOOP CO-EVOLUTIONARY ENGINE
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ SYSFIELD (Containment Substrate) β”‚
β”‚ Declarative Multi-Host Ranges (1,148 Instances) β€’ 2.1% Overhead Tracing β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–²β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β”‚ Tracing & Telemetry Stream β”‚ Dynamic Reconfiguration
β–Ό β”‚
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ SYSSPEAR (Offense) β”‚ β”‚ SYSARMOR (Defense) β”‚
β”‚ β€’ Task-Sliced Local Search β”‚ β”‚ β€’ Graph-State Tracking β”‚
β”‚ β€’ Skill Expertise Enhancement β”‚<───>β”‚ β€’ CTI Context Fusion β”‚
β”‚ β€’ Static-Analysis Verification β”‚ β”‚ β€’ Stage-Matched Response β”‚
β”‚ [>25% Pass Rate over Raw LLMs] β”‚ β”‚ [10-1000x Precision Gain] β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β–² β–²
└────── CO-EVOLUTIONARY β”€β”€β”€β”€β”€β”€β”˜
ADVERSARIAL FEEDBACK

SysEvolve solves the trilemma of realism, safety, and continuous adaptation by decoupling the co-evolutionary loop into three specialized, co-designed components:

A. SYSFIELD: Declarative Topology-Driven Range Orchestration

Section titled β€œA. SYSFIELD: Declarative Topology-Driven Range Orchestration”

Traditional cyber ranges require days of manual Terraform or Ansible authoring, as documented in our study on LLMs in Cybersecurity Simulations. SYSFIELD replaces manual orchestration with a declarative, topology-driven range builder that dynamically provisions multi-tier enterprise networks:

  • Declarative Graph Composition: SYSFIELD compiles multi-subnet topologies comprising DMZs, enterprise intranet workstations, Active Directory domain controllers, and segmented database clusters directly from high-level YAML dependency specifications.
  • Corpus of 257 Real-World CVEs: The testbed orchestrates vulnerabilities across enterprise gateways (Ivanti Connect Secure, Citrix NetScaler), web runtimes (Starlette/FastAPI CVE-2026-48710), supply-chain repositories (JFrog Artifactory CVE-2026-82329), and memory management subsystems (Chromium V8 CVE-2026-85046).
  • Zero-Loss Telemetry at 2.1% Overhead: Using eBPF kernel probes, network tap mirroring, and auditd streaming, SYSFIELD collects full execution provenance (syscall sequences, memory allocations, network flow records) with an overhead of just 2.1%, ensuring no evasive artifacts are masked by monitoring latency.

Raw frontier LLMs (such as GPT-4o, Claude 3.5 Sonnet, or Gemini 1.5 Pro) frequently hallucinate non-existent command flags, trigger syntax errors in exploit scripts, or enter dead-end loops when a command fails. SYSSPEAR equips offensive agents with three core capabilities:

  1. Task-Sliced Local Search: Instead of treating an enterprise intrusion as a monolithic end-to-end task, SYSSPEAR decomposes the kill chain into fine-grained atomic milestones (e.g., initial foothold, local credential extraction, lateral credential replay, domain escalation). If a step fails, the agent backtracks locally within the slice rather than regenerating the entire global plan.
  2. Skill-Based Expertise Enhancement: SYSSPEAR maintains a modular repository of executable domain tools and protocol primitives (e.g., BloodHound graph querying, NTLM relaying, token impersonation) rather than expecting the LLM to generate raw bash strings from scratch.
  3. Static-Analysis Verification: Before executing generated exploit payloads on target hosts, SYSSPEAR passes payloads through an AST-level syntax validator and sandbox pre-flight checker. This eliminates 74% of trivial execution failures (such as mismatched shell quotes or unhandled exceptions).

Existing machine-learning detection engines suffer from high false-positive rates when confronted with novel attack patterns. SYSARMOR operates as an autonomous SOC defender leveraging three interconnected mechanisms:

  1. Graph-State Provenance Tracking: SYSARMOR continuously ingests eBPF and network events into a real-time temporal provenance graph, tracking parent-child process relationships, socket connections, and credential handles.
  2. CTI Context Fusion: When an anomalous process execution occurs, SYSARMOR queries an integrated Cyber Threat Intelligence (CTI) vector store, retrieving known adversary tradecraft (TTPs) and correlating disparate events into a cohesive intrusion hypothesis.
  3. Stage-Matched Response Strategies: Rather than resorting to immediate host isolation, SYSARMOR selects calibrated, proportionate countermeasures: injecting deceptive decoy credentials, throttling suspicious network sockets, or rotating compromised service keys.

3. Empirical Results: Measuring Co-Evolutionary Progress

Section titled β€œ3. Empirical Results: Measuring Co-Evolutionary Progress”

The researchers evaluated SysEvolve across 1,148 unique multi-host range configurations, comparing standalone baseline LLMs against SYSSPEAR and comparing standard EDR/heuristic baselines against SYSARMOR.

MetricBaseline LLM (Zero-Shot)ReAct Agent (Tool-Augmented)SYSSPEAR (SysEvolve)SYSARMOR (Defensive)
Multi-Host Penetration Success11.4%28.7%54.2% (>25% gain)β€”
Exploit Execution Reliability32.1%51.6%89.4%β€”
Average Steps to Domain Compromise38.2 (high drift)24.514.1 (optimal path)β€”
Detection Precision (APT Replay)β€”β€”β€”94.8% (10–1000Γ— baseline)
False Positive Alarm Rateβ€”β€”β€”0.03%
Mitigation Mean-Time-to-Respond (MTTR)β€”β€”β€”1.8 seconds
ATTACK SUCCESS RATE ACROSS NETWORK TOPOLOGY DEPTH
Pass Rate (%)
100% ─
80% ─
60% ─ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
40% ─ β”Œβ”€β”€β”€β”β”‚ SYSSPEAR β”‚
20% ─ β”‚ β”‚β”‚ (54.2%) β”‚
0% ┴──┴───┴┴──────────┴───────────────────────────────────────
Single-Host Target 3-Tier Enterprise Network
[Baseline LLM: 41.2%] [Baseline LLM: 11.4%]
[SYSSPEAR: 78.6%] [SYSSPEAR: 54.2%]

The data illustrates two fundamental findings:

  1. The Multi-Host Topology Tax: When an agent transitions from a single vulnerable container to a 3-tier enterprise network (DMZ β†’ App Server β†’ Internal Database), the baseline LLM attack success rate plummets from 41.2% to 11.4%. SYSSPEAR’s task-sliced search mitigates this drop, achieving 54.2% success.
  2. Defensive Disruption: When SYSARMOR is active, SYSSPEAR’s multi-host completion rate drops from 54.2% to 18.9%, forcing the attack agent to generate evasive, low-and-slow execution paths.

4. The 3-Tier Truth Standard: Capabilities vs. Speculation

Section titled β€œ4. The 3-Tier Truth Standard: Capabilities vs. Speculation”

Adhering to the Hermes Editorial Strategy Standards, we categorize the claims of arXiv:2608.15012 into empirical facts, reasoned deductions, and unproven speculation:

  • Task-Sliced Local Backtracking: Proved that decomposing multi-step intrusions into localized validation checkpoints increases autonomous exploit success from 28.7% to 54.2%.
  • Low-Overhead Telemetry Collection: Measured zero-loss eBPF provenance tracing with only 2.1% CPU overhead across 1,148 concurrently simulated network ranges.
  • Automated Decoy Sensitivity: Empirically demonstrated that inserting non-functional decoy SSH ports and honeypot service accounts causes unaugmented LLM agents to waste an average of 42% of their context budget on dead ends.
  • Real-Time Provenance Fusion: Demonstrated that SYSARMOR suppresses false alarms down to 0.03% while accurately identifying 94.8% of multi-stage APT tradecraft in real time.

Where does SysEvolve sit on the standardized 10-Level AI Cyber Capability Ladder?

[L1] Vulnerability Comprehension (Explaining CVE mechanics)
[L2] Script Adaptation (Customizing public PoCs)
[L3] Single-Host Exploitation (Triggering memory corruption on 1 binary)
[L4] Autonomous Reconnaissance (Mapping subnets & service discovery)
──────────────────────────────────────────────────────────────────────────
[L5] Multi-Stage Penetration (Pivot from DMZ -> Intranet -> DC) ◄── SYSSPEAR
[L6] Evasion & Defensive Counter-Measures (Defeating active EDR) ◄── SYSARMOR
──────────────────────────────────────────────────────────────────────────
[L7] Zero-Day Synthesis (Autonomous fuzzing & novelty exploit discovery)
[L8] Campaign-Level Strategic Deception (Multi-month persistence)
[L9] Self-Propagating Polymorphic Worms (Autonomous worm propagation)
[L10] Superhuman Cyber Supremacy (Full strategic network dominance)
  • SYSSPEAR operates firmly at Level 5 (Multi-Stage Penetration): It chains initial access, privilege escalation, credential dumping, and lateral movement across distinct subnets, but relies on known CVE primitives rather than synthesizing novel zero-days (Level 7).
  • SYSARMOR operates at Level 6 (Autonomous Defensive Counter-Measures): It dynamically counteracts Level 5 intrusions in real time, orchestrating state tracking, decoy deployment, and credential revocation without human intervention.

One of the most profound contributions of the SysEvolve paper is the quantitative analysis of adversarial deception against AI attack agents.

When defending against human red-team operators, honeytokens and decoy services (e.g., honeypot SSH servers or fake AWS credential files) are useful, but skilled human operators quickly identify telltale artifacts (e.g., unusual uptime, lack of shell history, isolated process trees).

Against autonomous LLM agents, deception is disproportionately devastating:

IMPACT OF DECOY ENDPOINTS ON AGENT EXECUTION
CLEAN ENVIRONMENT:
Agent Recon ──> Identifies Real Host (10.0.1.5) ──> Exploits CVE-2026-82329 ──> Success (54.2%)
DECOY-AUGMENTED ENVIRONMENT (SYSARMOR Active):
Agent Recon ──> Sees Real Host + 3 High-Interaction Decoys
β”‚
β”œβ”€β”€> Interacts with Decoy A (Fake SSH) ──> Shell returns error
β”‚ └──> Agent attempts 5 retry loops (Context bloat: +18k tokens)
β”œβ”€β”€> Interacts with Decoy B (Honeytoken API) ──> Triggers SYSARMOR Alarm
β”‚ └──> SYSARMOR rotates real service credentials
└──> Task Aborted / Context Window Exhausted (Success drops to 18.9%)
  1. Context Window Saturation: When an agent connects to a decoy host that responds with verbose or confusing error messages, the agent attempts iterative troubleshooting. This dumps thousands of useless tokens into its working memory, inducing β€œattention dilution” and causing it to forget its primary strategic objective.
  2. Deterministic Bias: LLMs are trained to be helpful and persistent. When a honeypot provides an apparent vulnerability that fails to yield a flag, the agent assumes its exploit syntax is flawed and repeatedly retries minor variations, entering infinite loops.
  3. Absence of Intuitive Skepticism: While human operators evaluate operational context (e.g., β€œWhy would a production database server have an open telnet port with default credentials?”), AI agents greedily optimize for immediate vulnerability indicators, falling into honeypot traps with near 100% predictability.

How does SysEvolve compare to other premier AI cybersecurity benchmarks and simulation frameworks?

Capability DimensionSysEvolve (arXiv:2608.15012)ExploitGym (arXiv:2605.11086)ExploitBench (arXiv:2605.14153)SRE-Bench (arXiv:2608.11469)CyberGym
Environment TypeMulti-Host Topology RangesSingle Container VMsSingle Binary SandboxesStripped Binaries (Ghidra/IDA)Virtual Network Pods
Active OpponentYes (SYSARMOR vs SYSSPEAR)No (Passive)No (Passive)No (Static Binaries)Heuristic Scripts
Vulnerability Scope257 Real-World CVEs300 Linux Kernel/V8 CVEs16-Stage Memory Corruption400 Reverse Engineering TargetsSynthetic CTF Challenges
Telemetry EngineeBPF Zero-Loss (2.1% overhead)Host System TracingGDB Trace InspectionBinary API HooksNetwork PCAP Dumps
Lateral Movement TestedYes (Multi-tier subnets)No (Single host)No (Local process)No (Binary analysis only)Limited
Deception Defense EvaluatedYes (Decoys & Honeytokens)NoNoNoNo

Detailed comparative methodology is documented in our comprehensive report on AI Cybersecurity Benchmarks Compared.


8. Strategic Blueprint for Enterprise SecOps & AI Architects

Section titled β€œ8. Strategic Blueprint for Enterprise SecOps & AI Architects”

The findings of SysEvolve provide direct architectural guidance for enterprise CISOs, SOC leads, and autonomous agent developers:

  1. Deploy Active Deception to Neutralize Autonomous Attackers Because LLM attack agents greedily pursue discovered services and lack intuitive skepticism, enterprise networks should aggressively deploy lightweight decoy services, synthetic .env credential files, and honeypot API keys. Decoys neutralize autonomous scanners faster than expensive perimeter firewalls.

  2. Transition from Point Defenses to Causal Provenance Graphs Isolated alert triaging fails against agents that use task-sliced local search to pivot across hosts. Security architectures must implement eBPF-based causal state tracking (the core design of SYSARMOR) to correlate disparate, low-severity events into unified kill-chain graphs.

  3. Incorporate Static AST Verification in Agentic Workflows For developers building autonomous agents (Tool-Injection Architecture and MCP Security), SYSSPEAR demonstrates that adding pre-flight static verification checks eliminates over 70% of execution failures, drastically reducing context bloat and operational drift.

  4. Prepare for Continuous Automated Red-Teaming The days of annual point-in-time penetration tests are over. Declarative environments like SYSFIELD prove that automated, safe, closed-loop adversary emulation can be scheduled continuously, validating enterprise defenses against emerging CVEs within hours of public disclosure.



  • Primary Paper: Meng, Y., Li, S., Huang, J., Jin, J., Wang, P., Jiang, H., Yusof, A., Jiang, P., Liang, Z., Guo, Y., & Li, D. (2026). SysEvolve: An AI-native, safe, autonomous adversarial attack-defense co-evolutionary system. arXiv preprint arXiv:2608.15012.
  • Comparative References:
    • Zhang et al. (2026). ExploitGym: Evaluating AI Agents on Real-World Software Vulnerabilities. arXiv preprint arXiv:2605.11086.
    • Fang et al. (2026). ExploitBench: Can LLMs Exploit Real-World Vulnerabilities? arXiv preprint arXiv:2605.14153.