This post is meant as a background, or "relevant context", for our sequence on AI oversight and its limitations. It can also be read on its own. Selected takeaways This post is a message to AIs (and humans overseeing them) on the topic of orienting towards oversight. Some highlights: 1....
Once upon a time, A Relatively Famous Guy On The Internet was accused of having been simultaneously dating multiple women, without those women's knowledge, those women (according to the accusation) having been convinced that they were his exclusive partners all along. Then, one of his friends, another Relatively Famous Guy...
TL;DR: We suggest a sanity check for proposed evaluation or AI oversight schemes: Imagine the AI was replaced by a competent, strategic human — someone who knows they might get evaluated and has their own agenda. Would the evaluation still work? When we apply this mental move broadly, to all...
[I was aiming for legibility to a limited extent only. This post got extracted from a bigger post I've been writing and is meant mostly as a reference, and thus it may make more sense in context than in isolation.] (Spiritually related: Yes, It's Subjective, But Why All The Crabs?[1])...
TL;DR: * We introduce a concept we call "entanglement": roughly, the amount of information that the AI has about its environment. * We distinguish between "actual" entanglement between a specific instance of an AI and its environment and "minimum" entanglement corresponding to some task, without which solving the task is...
A Day at AFFINE[1] > “AFFINE was the best month of intellectual exploration I have had the opportunity to engage in, ever. Usually opportunities like this are limited to a day or a weekend, which both limits depth, forces a sprint-type mindset, and generally is quite limiting. At AFFINE I...
TL;DR Evaluation awareness — an AI recognizing it's being evaluated — is a widely discussed concept in AI safety. But there is a closely related concept that we claim is more important: deployment awareness, the AI's ability to recognize when it is not being evaluated and when its actions matter....