I think the characterization of OpenAI's injunction-focused model spec as inherently short-sighted or foolish is somewhat unfair; there are reasons to believe that injunction-heavy model specs generalize well, as counterintuitive as it feels. In particular, Anthropic released in May a report on "model spec midtraining," an approach which dovetails with the entanglement-first alignment approach you're detailing here. In that report, they empirically investigate the effects of adding reasoning and explanations to the model spec (the approach Anthropic seems ...
In section [2], you gesture at reward-instilled reflexes being instilled when simple, "cheap" tricks, not requiring much thought, are sufficient to trick graders and rubrics. By contrast, flexible reward-pursuit behavior should only be instilled when there are rewards to advanced techniques.
But if persona graders were so vulnerable to rhetorical flourish that they were saturated based on that flourish alone, then surely models would rarely engage with users at all except to deploy these tricks. Accordingly, the graders must have some additional engagement... (read more)