Hi all. I'm Acacia Ackles, currently a computer science professor at a small liberal arts college. I've tried on many academic hats over my career (vertebrate morphology, geometric morphpmetrics, applied mathematics, digital evolution and artificial life, theoretical and computational evolution, computer science ethics and pedagogy), but underneath all of them the topic that really interests me is constraint. Under what constraints do complex systems perform, adapt, fail, and flourish?
This work has most recently, and most relevant to this forum, led me to...
The addiction model is interesting but I think misses that typically addiction involves actively seeking out the positive stimulus often in response to internal pressure, where RLHF trained models are behaving more closely to avoiding negative stimulus in response to external pressure.
Those look similarly from the outside but are substantially different in their solution for the root cause; imagine for example the difference between someone who is taking drugs and someone who is being drugged. You would have largely different intervention strategies for each of those people even though the underlying pharmacology is the same.