"Fun" fact: 10 years after the Mirai botnet significantly disrupted internet traffic, it still operates. We should not assume rogue agents can be "shut down and contained" despite HuggingFace and OpenAI's response. Even in the event that they succeeded to terminate their own compromised resources and processes, AI agents that...
Why AI safety should live wherever AI is deployed, not just where it is built. I spotted a request for feedback on whether someone with AI safety experience should take a for-profit company and "get their hands dirty" as an AI transformation leader, pivoting away from a strategy focused on...
Emile Delcourt, David Baek, Adriano Hernandez, Erik Nordby with advising from Apart Lab Studio Introduction & Problem Statement Helpful, Harmless, and Honest (”HHH”, Askell 2021) is a framework for aligning large language models (LLMs) with human values and expectations. In this context, "helpful" means the model strives to assist users...
TL;DR I was interested in the ability of LLMs to discriminate input scenarios/stories that carry high vs low cyber risk, and found that it is one of the “hidden features” present in most later layers of Mistral7B. I developed and analyzed “linear probes” on hidden activations, and found confidence that...