There's a lot of discourse around how AI is getting really good at cybersecurity tasks. This is definitely true but the skills seem to be highly spiky.
For instance, there's strong evidence that frontier models are really good at discovering security vulnerabilities (Mythos), but they still seem to be weak at fixing security issues.
There was this interesting research paper released by 1password where they assessed Opus-4.8 and GPT-5.6's ability to write security patches. They found that only in ~1/4 of the cases did the model return an acceptable result, whereas in 3/4 of the cases it either didn't properly fix the vulnerability, changed features of the app, and in ~5% of the cases even introduced new vulnerabilities.
https://1password.com/blog/why-ai-generated-patches-still-require-human-review