Style prompts' effects on a small LLM's residual stream and outputs
Cross-posted from my blog, which has interactive figures. Epistemic status: exploratory. This analysis was conducted on one small open-weight model, gemma-2-2b-it, at temperature 0 with one run per prompt. I think these results are directionally correct, but I wouldn't bet on exact magnitudes. TL;DR: I added "style prompts" like avoid...
Sep 117