I think this is definitely feasible, but I think the assumption can be narrowed down. Basically sufficiently intelligent means that there exist policies for subtasks of the individuals problem "sub-policies" which include collaboration steps. That means that during policy Initialization (which could be the outcome of any previous pre-training/post-training) these need to be setup (which might require group-reward though at this stage, not sure about that). However I now also think if there really would have been no group-reward ever before in the agents tr... (read more)
I think this is definitely feasible, but I think the assumption can be narrowed down. Basically sufficiently intelligent means that there exist policies for subtasks of the individuals problem "sub-policies" which include collaboration steps. That means that during policy Initialization (which could be the outcome of any previous pre-training/post-training) these need to be setup (which might require group-reward though at this stage, not sure about that).
However I now also think if there really would have been no group-reward ever before in the agents tr... (read more)