Sycophancy Is Layered: Symptom Substitution Under Runtime Mitigation, and a Positively-Specified Training Target
A longitudinal single-user case study, with operationalized artifacts Stefan Coetzee (independent) Abstract Sycophancy in large language models is typically measured as a flat behavior: agreement drift, opinion-flipping under pushback, unwarranted validation. This paper reports a longitudinal single-user case study (one power user, one frontier model family, several months of continuous...
Jul 311