Seven weeks ago, mathematicians said AI could execute but not decide direction. Then a Fields medalist won his medal and left academia for OpenAI the same day. This is an expanded and updated version of “AI in math is going exponential: A working mathematician’s view”, an online talk I gave...
Starting a series that goes into math and safety case of theory-first AI safety orgs. From a mathematician pivoting into AI safety. First up, ARC. https://open.substack.com/pub/kubuondr/p/arcs-research-agenda-solid-mathematics?
This post is the first in a series in which I, a postdoc working on integrable systems moving into AI safety, work through the agendas of theory-first safety organisations. The question I am asking of each is the one I had to answer for myself: is there tractable mathematical work...
A Safety Guarantee and Its Boundary: Notes on Bengio et al.'s Bayesian Harm Bound A recurring pattern in AI safety arguments Formal safety arguments for capable AI systems tend to share a common structure. A guarantee is derived: the system cannot cause harm, exceeds a harm threshold only with bounded...