An Alien Mind
This is an unofficial automated linkpost. Continue reading at alignment.openai.com →
This is an unofficial automated linkpost. Continue reading at alignment.openai.com →
This is an unofficial automated linkpost. Continue reading at alignment.openai.com →
This is an unofficial automated linkpost. Continue reading at alignment.openai.com →
This time, I tried adding a bit more commentary to make things less dry. Preface * I show my discovery graph in (via …) blocks, those without usually come from my RSS reader or the algorithm of the site * This is approximately a 1 in 20 filter of content...
This is an unofficial automated linkpost. Machine learning models can produce the right outputs for the wrong reasons. Famous examples include a reinforcement learning agent that, rewarded for collecting a coin always placed at the right end of the level, learns to run rightward rather than to seek the coin...
Preface * I show my discovery graph in (via …) blocks, those without usually come from my RSS reader, or the algorithm of that site * This is approximately a 1 in 20 filter of content * This is very disorganized, but hopefully still useful. * Sometimes quotes are not...
This is an unofficial automated linkpost. We find that reinforcement learning on realistic scenarios targeting beneficial traits can produce broad improvements across dozens of benchmarks measuring aligned and beneficial behavior. These alignment gains generalize beyond the domains used for training and persist under adversarial pressure. As AI systems become more...