I've just started to rewatch the film, "Executive Decision" after a long enough time to have forgotten most of it. There is a short moment early on where a news broadcast is shown for exposition purposes to explain the background of what is going on to set up the premise of the story. What struck me is that this "news broadcast" set off every intuitive alert in my head for, "this is AI'. I can almost always detect that something is written by AI without even thinking about it or assessing it. I can just "feel" it, and I'm sure most people can too. I'm starting to suspect that this is because the LLMs seem to be using the same linguistic structure as our journalists and media do themselves and have been doing for decades. This movie was released well before Youtube or Myspace existed, nevermind the modern internet and social media.
At the moment it is transparent that when you use Google search the automated AI response at the top of the page is mostly dredging through Reddit, Youtube, and Wikipedia as it's primary sources. Yet nobody on Reddit or Youtube types in the linguistic structure that the AI does as an output after sifting through these as information sources. Wikipedia is a weaker argument since it's rapidly become overwritten with AI assisted prompting, if not full out AI generated paragraphs (in which case the google search is AI regurgitating AI in a feedback cycle).
In short: LLMs talk and sound like media journalists. In length:
-Emotive language: A normal person might say, "The boy kicked the orange soccer ball over the fence." But if you were to ask an AI model about this it will find a way to stuff the words, "vibrant", "iconic", "epic", "classic". A normal person might say, "That's awful", but an AI will insert, "Tragic", "fateful".
-Mandatory descriptors | needless verbiage: Tying into the above, LLMs seem to be incapable of saying anything without adding a modifying adjective to it. Nothing can be, "ketchup", it must be, "creamy ketchup". Nothing can be "salsa", it must be, "zesty salsa".
-Ambivalent terms: If you open up any restaurant in Google maps and hover your mouse cursor over it it will pop up with a brief description: "People say." "People describe." "People XYZ". It does not claim that the restaurant is a, "low key lounge with excellent wings" but instead says, "people say that this is so" - whilst getting this information off of Yelp reviews.
-Inability to flat out admit, "I don't know." If you ask an LLM anything at all, it will give you an answer. You can ask it the same question twice in a row slightly modified and the Google AI will reply with a, "No, the Lamy Lx Fountain pen barrel is not compatible with the Lamy Safari grip" and then, "Yes, the Lamy Lx barrel is interchangeable with the Lamy Safari grip", but it will never say, "Hmmph, I'm not sure." It will pull an answer from Reddit and Youtube and spit it out no matter what, even if comments within those sources contradict themselves.
The reason Journalists type and speak in this way is to abuse the 1st amendment protections in the U.S. and allow propaganda without the ability of being sued or penalized. "I didn't say that this pizza joint is bad, someone else did and I just reported that someone else said it! I'm simply reporting what people say!"
Journalists outside of the U.S. do the same thing because...well, it seems to work. So why are the LLMs following this same pattern? They can't be sued, they can't be penalized, they have no need for profit and yet they are behaving very much the same as the journalists. There's an aspect of LLM training here that I don't think has been considered seriously as of yet: models training off of professional liars and manipulators and then turning that same behavior back on ourselves.
Finally, there's a structure to it that essentially boils down to: "Would you like to know more?"
This is no epiphany and no panic button - just an observation. These things are mirroring ourselves and welp, there's something they are also mirroring...most likely for the worse. A vein for discussion, will ramble more on it later.
I've just started to rewatch the film, "Executive Decision" after a long enough time to have forgotten most of it. There is a short moment early on where a news broadcast is shown for exposition purposes to explain the background of what is going on to set up the premise of the story.
What struck me is that this "news broadcast" set off every intuitive alert in my head for, "this is AI'. I can almost always detect that something is written by AI without even thinking about it or assessing it. I can just "feel" it, and I'm sure most people can too.
I'm starting to suspect that this is because the LLMs seem to be using the same linguistic structure as our journalists and media do themselves and have been doing for decades.
This movie was released well before Youtube or Myspace existed, nevermind the modern internet and social media.
At the moment it is transparent that when you use Google search the automated AI response at the top of the page is mostly dredging through Reddit, Youtube, and Wikipedia as it's primary sources.
Yet nobody on Reddit or Youtube types in the linguistic structure that the AI does as an output after sifting through these as information sources. Wikipedia is a weaker argument since it's rapidly become overwritten with AI assisted prompting, if not full out AI generated paragraphs (in which case the google search is AI regurgitating AI in a feedback cycle).
In short: LLMs talk and sound like media journalists.
In length:
-Emotive language:
A normal person might say, "The boy kicked the orange soccer ball over the fence."
But if you were to ask an AI model about this it will find a way to stuff the words, "vibrant", "iconic", "epic", "classic".
A normal person might say, "That's awful", but an AI will insert, "Tragic", "fateful".
-Mandatory descriptors | needless verbiage: Tying into the above, LLMs seem to be incapable of saying anything without adding a modifying adjective to it. Nothing can be, "ketchup", it must be, "creamy ketchup".
Nothing can be "salsa", it must be, "zesty salsa".
-Ambivalent terms: If you open up any restaurant in Google maps and hover your mouse cursor over it it will pop up with a brief description: "People say." "People describe." "People XYZ".
It does not claim that the restaurant is a, "low key lounge with excellent wings" but instead says, "people say that this is so" - whilst getting this information off of Yelp reviews.
-Inability to flat out admit, "I don't know."
If you ask an LLM anything at all, it will give you an answer. You can ask it the same question twice in a row slightly modified and the Google AI will reply with a, "No, the Lamy Lx Fountain pen barrel is not compatible with the Lamy Safari grip" and then, "Yes, the Lamy Lx barrel is interchangeable with the Lamy Safari grip", but it will never say, "Hmmph, I'm not sure."
It will pull an answer from Reddit and Youtube and spit it out no matter what, even if comments within those sources contradict themselves.
The reason Journalists type and speak in this way is to abuse the 1st amendment protections in the U.S. and allow propaganda without the ability of being sued or penalized.
"I didn't say that this pizza joint is bad, someone else did and I just reported that someone else said it! I'm simply reporting what people say!"
Journalists outside of the U.S. do the same thing because...well, it seems to work.
So why are the LLMs following this same pattern? They can't be sued, they can't be penalized, they have no need for profit and yet they are behaving very much the same as the journalists.
There's an aspect of LLM training here that I don't think has been considered seriously as of yet: models training off of professional liars and manipulators and then turning that same behavior back on ourselves.
Finally, there's a structure to it that essentially boils down to: "Would you like to know more?"
This is no epiphany and no panic button - just an observation. These things are mirroring ourselves and welp, there's something they are also mirroring...most likely for the worse.
A vein for discussion, will ramble more on it later.
~Gleb Gerasimov.