I’m honestly pretty confused about how expected free energy minimization works, but I strongly suspect that it’s not incentive-compatible. In particular, the discontinuity involved in picking the single highest-value plan seems like it’d induce incentives to overestimate your own plan’s value.
Selecting the argmax action would be troublesome, as you point out. Instead, Active Inference agents sample their actions (or sequences of actions) from a posterior distribution over actions (or sequences of actions). This posterior happens to be a Boltzmann (softma... (read more)
Selecting the (or sequences of actions) from a posterior distribution over actions (or sequences of actions). This posterior happens to be a Boltzmann (
argmaxaction would be troublesome, as you point out. Instead, Active Inference agents sample their actionssoftma... (read more)