What makes an action good?
Good outcomes and good actions
What does it even mean to take a ‘good’ action?
Let’s formalise how action-taking generally looks within a simulation of the outcomes. It will be useful to separate the two into different partitions.
Note that the downstream outcome iteration can be replaced by any kind of downstream simulation that is connected to the action iteration; we only think of it as a single iteration here for convenience.
From the probabilistic perspective, we only ever see one action-to-outcome pairing in most real world action-taking scenarios. The fuller range of possibilities that a simulation represents is much larger.
Given that only a single path through the range of possibilities is actually taken by the real world, it’s quite possible in many situations that any form of data can be misleading when relating actions to their associated outcomes.
Put another way, if there is only one result to validate an outcome from a given action: how do we know we didn’t just ‘get lucky’ at that point in time?
The difference between luck and skill
… Add stuff here about using the simulation to validate luck vs skill … Note that this depends on the model … Conclude that all of this relies on the simulation being ‘correct’