Multi-Agent Coordination Lets Us Pour More Compute Into Post-Training
Training models to coordinate across multiple agents gives us a way to productively spend substantially more compute during post-training than current single-agent RL setups. From the recent Dwarkesh Podcast with Noam Brown: > Noam Brown > > But that’s one data point. We don’t know how long it would take...
Sep 227