What is a policy?

Defining it

Action-taking policies define the logic which take in the state of the world and map it to a taken action at any moment in time.

Note that the downstream iteration of simulated ‘outcomes’ from a given action can be replaced by any kind of downstream simulation that is connected to the action iteration; we only think of it as a single ‘iterate simulation partition’ here for convenience.

Sensitivity and generalisation

Which action-taking policy logic finds the best actions?

How sensitive is this choice to changes in the data? How sensitive is this choice to changes in the outcome model?

Answers to both of these questions tell us how the action-taking policy generalises within the problem domain to alternative scenarios and underlying system mechanisms.

How sensitive is this choice to changes in the action parameters?

In much the same way as it does in learning simulations of the real world answering this question tells us how sloppy the action-taking policy is, and as a result how generalisable it could be to other problem domains.