This post follows The iVAIS Manifesto: Safety Through Character, Not Compliance. A natural objection to the approach of iVAIS suggested earlier[1] may be that it just replaces one proxy with another. If ordinary alignment approaches fail because models learn to optimize proxy goals rather than genuinely acquiring the intended target,...
Masaharu Mizumoto, Mads Udengaard, Rujuta Karekar, Mayank Goel, Daan Henselmans, Nurshafira Noh, Saptadip Saha, Pranshul Bohra TL;DR * AI safety requires a shift to character-based, virtue-centered alignment * Rule-based and principle-based alignment are conceptually insufficient * Constitutional AI is still action-based ethics * Mechanistic interpretability cannot ensure control * The...
TL;DR: Constitutional AI remains largely rule-based rather than fully character-based, as it should be. We propose a virtue-ethical alternative based on holistic human intuitions. Introduction Anthropic’s Constitutional AI proposes an ambitious strategy for aligning advanced AI systems. Instead of relying solely on human feedback, the model is trained to follow...