bling
Message
@blingdivinity on twitter
you may know me there from my AI takes, leaking llm CoTs, and asking llms about nonexistent seahorse emojis
2
the crowd attack examples you give nearly guarantee that the terrorists will be caught or killed when carrying out the attack. in comparison, you can operate drones from a distance and with way less risk. spending a few thousand dollars on drones is many times cheaper compared to one of your men’s lives. this means very small groups could continue doing these attacks many times before they get caught by the state. and if they’re competent enough, they may get away with it completely.
yes, quality data is much more abundant in easy to verify domains for the reasons you have spelled out. this is solid evidence for verifiability not just being essential for rl, but for imitative learning as well. so your post title is misleading since your core claims are still about verifiability, just verifiability on the side of the humans that created and curated the pretraining data and not only on the side of the rlvr graders verifying correct answers.