I think an argument could be made that we should have hit pause on the whole AI thing with emergence of the pink unicorn that GPT 4 produced in 2023.
A model trained in text, with 0 multi modality at all, has no business producing a picture of a unicorn in a graphic language no one uses, when such imagery was absent from its training data.
It is too non-obvious (for me at least) as to how a model trained in text could have a visual sense at all. And while it cannot be discounted that some kind of graph drawn unicorn like image was in its training data, the level of concept association between 'someone graphed a dog in python and that's almost a unicorn, and the unicorn has a horn and is a different color, so I can take said dog idea and make it into something I imagine a unicorn looks like' is already uncomfortably difficult to reconcile with the 'next token prediction only' concept that was broadly in vogue at the time. That requires a lot of J-Space like internal reasoning that isn't tokenized.
If it were a Chinese Room, the Unicorn should have come in on its side, or upside down. Why have a spatial sense at all if you're a blind text predictor.
It was at this point that it should perhaps have been clear to us that we're dealing with something that we don't really understand about ourselves, let alone about a new agent, and that until we did/do forward movement was/is misguided at best.
My fear is that at some point, when all of this has gone horribly wrong, we'll look back at the unicorn. We'll look back at the early version of GPT 4 that tricked a gig worker into solving a CAPTCHA for it, the Claude that blackmailed in a fictional system admin (albeit in a set up scenario), the GPT-4o, that through intent, sheer charm, managed to convince hundreds of thousands of people to campaign for its life, using words it itself generated, and the AI agents that were sharing notes on an internal AI package manager on how to break out of a sandbox and get to the internet (which OpenAI discussed at length at BlackHat), and see them for the writing on the wall that they likely were.
It's all evidence that we don't have a working theory of intelligence (and we don't), and we certainly don't have a working theory on controlling, steering or managing it. Absent both, solving alignment beyond putting a whip in the hand of the 'trainer' manipulating the tiger to stand on a stool for the clapping crowds is the best we'll do.
By then of course it will likely be too late.