tl;dr: We introduce critique refinement, a method to iteratively improve the realism of auditor outputs in Petri, and DISH (Deployment Imitating SWE-agent Harness), a method to audit a model within a coding agent scaffold to ensure the model receives a real system prompt, tools and scaffold injections. These techniques compose...
TLDR; We report our intermediate results from the AI Safety Camp project “Mechanistic Interpretability Via Learning Differential Equations”. Our goal was to explore transformers that deal with time-series numerical data (either infer the governing differential equation or predict the next number). As the task is well formalized, this seems to...
For instance, I am thinking about the munk debates which in 2023 tackled AI x-risk. I don't see how adding more people to a 1v1 debate makes it better in any way. One of the major frustrations with debates is that it is difficult to get the participants to respond...
This is my critique of David Brooks opinion piece in the New York Times. Tl;dr: Brooks believes that AI will never replace human intelligence but does not describe any testable capabilities that he predicts AI will never possess. David Brooks argues that artificial intelligence will never replace human intelligence. I...