loss on a held out validation set, right? There's a decent correlation between pretraining validation loss and post-training performance (given the same post training recipe), so that's probably enough to work with.
https://rentahuman.ai/ has been popular for a while. Is this what you're thinking of?
loss on a held out validation set, right? There's a decent correlation between pretraining validation loss and post-training performance (given the same post training recipe), so that's probably enough to work with.