Moreover, I want AI safety solutions that produce explicit, quantitative safety guarantees that are underpinned and motivated by explicit, auditable assumptions. I don’t think that purely empirical methods are adequate for producing safety assurances that are satisfactory or acceptable for very powerful AI systems.
I think this is the only sane approach. I have so far not been convinced that empirical methods will lead anywhere, since they only evaluate the models behavior given an environment but they make no claim or assurance about how the model will act outside of that.
I think this is the only sane approach. I have so far not been convinced that empirical methods will lead anywhere, since they only evaluate the models behavior given an environment but they make no claim or assurance about how the model will act outside of that.