I may be misunderstanding the intended scope of the conjecture in part c of question 2, but I think there is a simple counterexample as currently stated. Two reward functions could agree on all actions that are ever optimal, while ranking two permanently suboptimal actions differently. They would then have the same optimal policies for every transition function, but not the same full policy ordering.
This seems slightly different from the stated exception, since the optimal policy could still depend on the transition function. Perhaps the conjecture needs to exclude policies that can never be optimal?
I may be misunderstanding the intended scope of the conjecture in part c of question 2, but I think there is a simple counterexample as currently stated. Two reward functions could agree on all actions that are ever optimal, while ranking two permanently suboptimal actions differently. They would then have the same optimal policies for every transition function, but not the same full policy ordering.
This seems slightly different from the stated exception, since the optimal policy could still depend on the transition function. Perhaps the conjecture needs to exclude policies that can never be optimal?