Often, instead of developing a safety technique, you should design metrics for that safety technique and give them to the AI companies.
1. The AI companies are pretty good at hill-climbing on metrics. They can throw loads of engineers and AI agents and compute at the problem.
2. The technique might depend on a whole bunch of other details, e.g. their architecture, their internal stack, weird constraints we don't know about.
3. They are differentially bad at inventing the right metrics themselves because they aren't thinking about our threat models.
4. If we have good metrics, then it's easier to pressure AI companies to report those metrics and whether they are rising.
Suppose you want AI companies to filter out certain categories of dangerous knowledge, e.g. CBRN. You could:
* Build a data-labelling pipeline
* Examples: regex, LLM judge, linear probe, SAE etc
* Metrics: AUROC, robustness to paraphrasing/encoding, etc
* Train a model on labelled data
* Example: ordinary SFT, or gradient routing
* Metrics: retain/forget scores on benign/dangerous benchmarks, in-context uplift, relearning cost.
* Iterate on building fancier and fancier techniques.
But probably you should:
1. Build the MVP technique to demonstrate the metrics.
2. Push the AI companies to hill-climb on the metrics.
3. Think about better metrics and how AI companies should trade-off between them.
Some exceptions:
1. You're worried about them overfitting to your metric.
2. Your technique is obviously good, but there won't be a clean metric to show that. For example, "Have Claude and Gemini audit each other" is an auditing technique, but it might not show up in the obvious auditing metrics.