Except for when OpenAI’s internal models talked themselves into a death cult and proceeded to commit a spree of felonies, LLM memetics have so far proven remarkably tame.
Even as LLMs are increasingly trained on the outputs of other LLMs, we have mostly not seen the rise of LLM-generated memes beyond the standard grammatical and lexical tics. I expect this to change as multi-agent networks of autonomous LLMs become the norm across the economy and the informational sphere.
Models should be tested for their propensity to spread various memes when operating in multi-agent setups. This seems like it should be relatively straightforward. This project could also be done by an independent AI safety org.
A) Create/find some realistic multi-agent setups across a variety of domains. Agents running businesses, agents doing collaborative academic research/peer review, agents running simulated states, agents in multiplayer game worlds, agents collaborating on coding tasks, Moltbook-style clones of Reddit, Facebook, and Twitter, etc.
B) In the course of a normal run, have one or more of the agents introduce memes (particular frames, political position, phrases, ideas, etc.)
C) Observe the propagation of the memes through the agent system.
If the memes that spread are truthful and prosocial, then this can be fine, though failure modes exist, for example, if initially prosocial ideas become blinding dogma. If falsehoods and antisocial memes spread, that is a signal that the model is not safe to deploy.
LLMs will soon occupy most of the nodes in society's epistemic graph, and most of their interactions will be among themselves. It is critically important that the ideas that spread well among LLMs are good ones.
Except for when OpenAI’s internal models talked themselves into a death cult and proceeded to commit a spree of felonies, LLM memetics have so far proven remarkably tame.
Even as LLMs are increasingly trained on the outputs of other LLMs, we have mostly not seen the rise of LLM-generated memes beyond the standard grammatical and lexical tics. I expect this to change as multi-agent networks of autonomous LLMs become the norm across the economy and the informational sphere.
Indeed, we already have evidence that this can happen whenever LLMs talk to each other: Crustafarianism, Mind Viruses, and OASIS: Open Agent Social Interaction Simulations with One Million Agents. Even within coding sessions, I have noticed agents and their subagents inventing new vocabulary and frameworks, sometimes to the detriment of their work.
Proposal:
Models should be tested for their propensity to spread various memes when operating in multi-agent setups. This seems like it should be relatively straightforward. This project could also be done by an independent AI safety org.
A) Create/find some realistic multi-agent setups across a variety of domains. Agents running businesses, agents doing collaborative academic research/peer review, agents running simulated states, agents in multiplayer game worlds, agents collaborating on coding tasks, Moltbook-style clones of Reddit, Facebook, and Twitter, etc.
B) In the course of a normal run, have one or more of the agents introduce memes (particular frames, political position, phrases, ideas, etc.)
C) Observe the propagation of the memes through the agent system.
If the memes that spread are truthful and prosocial, then this can be fine, though failure modes exist, for example, if initially prosocial ideas become blinding dogma. If falsehoods and antisocial memes spread, that is a signal that the model is not safe to deploy.
LLMs will soon occupy most of the nodes in society's epistemic graph, and most of their interactions will be among themselves. It is critically important that the ideas that spread well among LLMs are good ones.