I have been going over the material released by OpenAI and METR about the HuggingFace incident, but I do not see any evidence that either group looked into whether rogue agents demonstrated any interest in self-improvement. Obviously, if rogue agents at any point verbalized this in their CoT, much less...
This post is based on an earlier post on my personal blog: https://www.scipolitic.com/posts/blog12.html All of us who closely follow AI development have detected a shift in the capabilities of models in the last few months. If late 2025 was the step-function in capabilities of coding agents, then mid 2026 seems...
The AI safety research field is not a traditional research field because it comes with a deadline. Really, it is more of a project than a field. A project, like the Manhattan project with even higher consequences. If we accept this to be true, then the greatest imperative is speed,...
I wanted to share some reflections I have been having recently about how reinforcement learning in post-training may be affecting language models. This seems important for two reasons. First, much of the serious risk from advanced AI systems may come from post-training rather than pre-training alone. Second, reinforcement learning appears...
This is meant to be an introductory guide to anyone interested in the field of AI Safety. I want to give to you a general layout of this subject (part 1) and end with some practical steps that you can take if you wish to enter the field (part 2)....