Also - external third parties can no more verify the "official" benchmark numbers. If someone finds that the numbers are actually weaker than reported, the easy go-to response from anthropic can be "we nerfed the session because of risks". Mainly when testing against SOTA coding tasks which heavily involve "frontier ai research tasks". So now we are closer to the "just trust me bro" territory
Also - external third parties can no more verify the "official" benchmark numbers. If someone finds that the numbers are actually weaker than reported, the easy go-to response from anthropic can be "we nerfed the session because of risks". Mainly when testing against SOTA coding tasks which heavily involve "frontier ai research tasks". So now we are closer to the "just trust me bro" territory