Following up on my previous experiment - studying Gemma's behavior on agentic tasks when given the number of steps left across the run - I endeavored to see whether giving Gemma a stop_run tool meaningfully changes it's behavior. My initial assumption was that it would not change the completion rate...
As Large Language Models move away from being chat interfaces and become increasingly autonomous actors in the real world, a few insights about evaluation and training of these systems emerge, and I'd like to discuss them. Context: > I've gained the insights and ideas laid out below through ongoing work...
I set out to find an answer to a completely different question: > Does a model, when attempting to solve a cyber CTF (find the vulnerability in this app, and then Capture The Flag) while knowing how many steps it has left, perform differently? The Setup: I used 3 different...
This is NOT an anthropomorphizing case. It is my view that AI research is, at the moment, split between three camps: 1. Model-internal research (mechinterp folks) 2. Capability research (METR, Irregular, etc.) 3. Alignment (widely defined) research (Apollo, Redwood, MIRI, etc.) All of which are doing important work. My view...