Epistemic Status: Exploratory synthesis. Background in mathematics/statistics (UChicago) and principal-agent problems (UT Austin doctoral ABD), with extensive study of psychotherapy literature motivated by personal research into consciousness variation. New to ML implementation details but confident in the conceptual mappings. Seeking technical feedback on proposed experiments.
Tl;dr: I suggest therapeutic techniques from a variety of psychotherapeutic schools of thought can inspire new approaches to AI learning and alignment. I reinterpret three recent AI/ML papers in the language of psychotherapy and propose three testable training methods inspired by common psychotherapeutic interventions.
Introduction
I've been meaning to post this essay for a while, and yesterday's top paper on Hugging Face, by Cui et al., finally convinced me to do it. Their paper provides a timely opportunity to map the language used by ML and AI engineers to the language used by humanistic psychotherapists—a translation which is more important now than ever as we struggle with increasingly stubborn problems in AI alignment, while simultaneously developing AIs whose capabilities are rapidly superseding those of humans.
I'll provide a high-level overview of my understanding of the paper and map it back to ideas from humanistic psychotherapy. I will then consider a few related papers which tie nicely to psychotherapeutic principles, and end with a few proposals for experiments. I am new to AI alignment, welfare, and interpretability research and I look forward to comments which can help me deepen and clarify my inevitably imperfect understanding of the papers I am citing.
The Core Analogy: Policy Entropy as Behavioral Flexibility
The Cui et al. paper "aims to overcome a major obstacle in scaling RL for reasoning with LLMs, namely the collapse of policy entropy."
Think of "policy" as the individual in therapy. The individual has a behavioral repertoire—a probability distribution of potent