Top postsTop post
TL;DR TypeSafe AI has introduced Jev - a new class of frontier model trained to make fast, structured decisions, rather than generating free-form text like a chatbot. It takes unstructured state as input and returns type-safe, structured outputs with confidence scores. I aim to use Jev as the trusted monitor...
TL;DR Obfuscated activations are internal model states that have been adversarially optimized to appear benign. They evade activation-based detectors, and still result in harmful model behaviour/output. Obfuscated Adversarial Training (Bailey et al. (2024)) is a method to train the model (not the monitor) to preserve monitor-detectable harmfulness representations even under...