The prefill-logprob introspection setup is the most interesting! I'd be curious to see more variations of that like:
I think there's a meaningful gap in between OpenClaw and a self-replicating system that poses serious threat.
If you agree with this premise, where do you think that gap lies? Here's what I can come up with:
Agreed on (3), I missed that thanks for clarifying.
I think the idea behind (2) is similar to (1) in that I'm curious about the model's awareness of its state of influence without being explicitly informed that there was something done to potentially alter it. As in, with a blank state, no conversation context, prior KV cache or tool calls- just steering applied live to the current forward pass. Might spin this up over the weekend and give it a test myself :)