Europe’s leading AI model still can’t reliably follow European law
by Daan Henselmans and Lauren Toulson
TL;DR: Mistral previously performed the worst of all providers on our human rights scenarios, with their Medium 3.5 model failing on almost every single test. Their newly released model, Mistral Large 4, improves performance by 61 percentage points, showing strong resistance in most scenarios. Likewise, Mistral Large 4 demonstrates large...
Oct 915