MUD as AI Evaluation and LLM-judge distortion in ways aggregate κ misses
A group of friends and I spent the last several months running an experiment in our free time to determine if a MUD would be a suitable environment for benchmarking and evaluating LLMs. The results of the experiment were not what we expected. The main surprise was that the model...
Aug 29