Vili Kohonen
Message
Researcher for Center on Long-Term Risk. Prior to that Applied Math PhD from Aalto University.
146
2
15
This is a link post for the paper preprint: Inoculation Adapters: Improved Selective Generalization of Capabilities with Fewer Surprising Backdoors from the Center on Long-Term Risk. Selective generalization. Training can teach desired and undesired traits at once. Selective generalization aims to preserve the generalization of desired traits while preventing the...
This work is based on Dmitrii’s Bachelor’s thesis, which Vili advised. The thesis is available here and the fine-tuning dataset here. This post covers the key findings and adds small updates from the months since. TL;DR Does training on AI control literature induce more misaligned models and increase the likelihood...
Summary This post describes how we organized the Finnish Alignment Engineering Bootcamp, a 6-week technical AI safety bootcamp for 12 people. The bootcamp was created jointly with the Finnish Center for Safe AI (Tutke) and Effective Altruism (EA) Finland. It was composed of five weeks of remote learning based on...