Top postsTop post
Shunk
Message
New to AI safety, writing my thoughts on research papers as a way to learn and communicate with others in the space.
Substack (same posts): https://substack.com/@shunks
22
5
8
Notes pt. 4, decided on MoReBench! Paper Link: https://arxiv.org/abs/2510.16380 Summary (What) AI systems are being increasingly used for decision making, but how are models actually making their decisions? MoReBench has 1000 moral scenarios, each with a rubric made by experts. MoREBench was created for both AI as an advisor and...
Not a very original paper review, so here are my concise notes on the topic, and near the bottom some of my thoughts on future experiments that I think would be interesting. Part 3 of notes series! Important note: This post was written based on the Anthropic blog post, and...
Part two of my notes series, with hopefully many more to come. Comments are much appreciated. Article: https://www.anthropic.com/research/multiagent-systems Paper summary (What) Institutions are designed for people, but soon, we'll have far more agents interacting in the real world. Agent-to-agent interactions could become the most common form of interaction before we...
Hey everyone, trying out taking notes on various AI safety papers as a method of increasing my knowledge and actually remembering the notes I take. I'm pretty new to the field, and I think that in the past, some of my notes have been too similar to the original paper...