Neuro-symbolic RL agent that learns to pentest networks it has never seen — GPT-4o compiles CVE preconditions into a Z3 action mask over a GraphSAGE PPO policy. Zero-shot attack-graph transfer, negatives disclosed.
research reinforcement-learning cybersecurity penetration-testing smt-solver autonomous-agents ppo graphsage attack-graph graph-neural-network neuro-symbolic z3-solver llm nasim action-masking netwrok-attack-simulation
-
Updated
Aug 7, 2026 - Python