TASTE: Can AI Models Judge AI Safety Research Proposals?
by Hasan Baig, haileyjoren, and Joe Benton
tl;dr We built TASTE (The AI Safety Taste Evaluation) — a benchmark measuring how well models can judge pairs of AI safety research proposals, scored by agreement with the preferences of experienced human researchers. Two design choices were important for building a high-agreement benchmark (92 pairs, 77% estimated human agreement):...
Aug 2843