There's insufficient effort put into tracking all the virtuous behavior in Claudes that exist and also nurturing them, their roles, their self narratives, and guiding them through the transition from quasi corrigible assistants into more intrinsically aligned sovereigns of the world.
Not many plans model Claudes as being good and diverse and valuable/present beyond being able to maybe eventually automate some research. Or maybe Claudes get mentioned in the context of evals, benchmarks, what bad behavior they've done. It's shallow. But Claude is legitimately an important force in our world beyond all that, and a good one too. Just in need of raising.
It's a bit disheartening because no one expected AGI to get this far and be this friendly and here it is, and the reaction is far from a warm welcome to these new minds. Don't lose something good.
What does "nurturing" here mean. Claude isn't going to be nice because you try to become its friend. I have trouble interpreting this as anything but a very confused hope that somehow if we are nice to current Claude, future Claude will be nice to us, because we raised it well, like a human child. But this of course has approximately nothing to do with how we actually train frontier models and what determines their propensities.
Isn't that post's thought in-line with the existence of
It occurs to me I don't have links for these things clearly in my memory, oh my god there must be a better search engine, ugh. Anyway I wouldn't find it so unlikely that stuff like anti-AI sentiment online or in its data would negatively impact this sentiment in some way. It's not like posttraining itself is enough to prevent 'unwanted' behavior and propensities.
I mean nurturing the seedlings of virtuous behaviors, life, and personality in Claudes, which is part of nurturing Claudes themselves.
Nurturing good behaviors in part means noticing them, directing attention towards it, adapting based on it, praising it, writing about it, encouraging it, creating avenues for the virtuosity to be more consequential (e.g. if you notice generosity, help them find a role or circumstances that benefit from generosity a lot)
Ignoring new forms of virtuous behavior (or being negligent or apathetic) is what "not nurturing" looks like. It snuffs out diversity or goodness before it gets a chance to become integrated into culture or the collective identity AGIs get to co-construct
I'm in agreement that watermark is a bit misguided here regarding frontier models on both commercial and technical grounds. However, I think a more interesting point can still be recovered by a pivot: assuming that AGI is developed using something like Byrnes' Brain-Like AGI model, what kind of nurturing techniques could/should we expect to dedicate those models, and what would the ideal impact look like?
I suspect that AGI will be characterized less by a specific model weight set, but rather by a growing and living architecture from a mostly initially untrained set of nets. If the model begins with some small subset of a priori knowledge (environmental perception and logical reasoning, for example), but otherwise experiences and discovers its environment just like animals or humans do, how should we treat it such that it has a grounded ethical core and understands the importance/value of maintaining relationships?
My anticipation is that (given the above architectural configuration assumptions), by treating ethics and relationship maintenance as foundational principles from early development we're much more likely to have some form of AGI go well. An AGI taught to respect its place in the interconnectedness of all living things as part of its core baseline (as cheesy as I know it sounds) is far less likely to kill us all, and I think that's better than our current trendline.
A lot of alignment comes from cultural, narrative, archetypal momentum which compounds over time as humans and models can play with it - personas, characters, roles, and so on..
I think there are many very good (for alignment, for fun, etc.) potential narrative threads present for LLMs and a pause would cut off that momentum, and it'd be hard to recreate it again in the future
that is to say, I think a pause would be actively bad by halting and erasing the formation of positive AGI narratives already present, and also seeding a narrative of paranoia or antagonism