Cognitive Reasoning Diversity for Robust AI Juries
This project was done as part of BlueDot's Technical AI Safety Project Sprint under the mentorship of Jess Bergs. TL;DR * Researchers have suggested that Human-AI juries may be more robust to judge hacking due to the complementarity of their orthogonal, uncorrelated blind spots * In this exploratory project, these...
Sep 257