I am not aware of the provenance of this particular misunderstanding, but this comes to mind as the kind of thing that could result from the tactic that you warn about: https://x.com/search?q=andrew yang polluted&src=typed_query
This might have been the post that gave me this idea. I remember being shocked and skeptical when I first saw the claim being made.
To me, Yang's claim is extraordinary and doesn't make sense in a few different ways. The attack I warn about in this post will (I think) be:
1) coordinated: stemming from multiple accounts at a similar time, possibly including news outlet/s that have built trust in the community but is funded by an adversary (MTS?)
2) Warning signs of information being false for a true memetic attack will feel much weaker than Yang's claim, you will have a harder time spotting them. However the pull the share the information will be just as strong as the pull I originally felt to share Yang's claim. I imaging myself thinking something like: "omg, another warning shot! And no lives were taken, no harm caused to a creature. Excellent! This evidence makes our claims stronger!"
Still, I think the Yang thing is a good example that fits partly into the type of Twitter post I would expect to see.
Addendum: A friend on twitter also points out that, likely the people who see this and my twitter post are not the ones who need warning about an attack like this. Probably the attack focuses on the movement broadly. For example, it'd really be bad if policy makers attempting to regulate ASI were seen reposting the misinformation. Though, if I were running an attack of this sort, I'd really be hoping catch some big fish too (Nate, Eliezer, Daniel, Jacob, Jeffery, etc).
Most likely current candidate is actually Andrew Yang's confused interview where he talked about "self-replicating code" and AI companies needing to create their own "synthetic internet" in order to prevent data poisoning while training their models or whatever. As far as I can tell, this was an exageration of Tristan Harris's exagerated claims on Glenn Beck.
I really hope that Tristan is making a good-faith mistake, but he clearly is merging OpenAI search agents leaving communications all over the internet (completely true) with spontaneous prompt injection in gpt COT summaries (also completely true) to form some sort of "AIs are leaving prompt injections on the open internet to poison future training runs" (completely wrong, and he should be publishing immediately if not).
If intentional, this counts as a clear lie about technical details. It would be extremely bad if this is where Yang got his ideas from. Glenn Beck is one of the most technically literate people in political punditry - exploiting his trust cynically is not going to end well and is unethical besides.
While this is relevant to my work at MIRI, I have not checked these ideas with anyone else on the team and am posting this on my personal LW account. These views are my own[1]. And to be honest, I am writing this mostly to remind myself of my weakness.
---
I expect one (or many) adversarial memetic attacks aiming to trip you up, perhaps consisting of fake leaks relating to dangerous stuff happening in the labs. Specifically, worrying incidents that may fit snugly within your worldview, leaking from multiple sources including news outlet/s, but not confirmed/confirmable by a primary source. Think rumors about exfiltrated weights, AIs attempting to create viruses, agent swarms hacking into and gathering information from nuclear infrastructure, etc.
An easy way to remove status from a movement is to trip it up: make it fall for a misinformation trap in public, then use that slip-up to discredit the movement for all time. The game is set for a memetic attack like this. There's a well-resourced group waiting for your screw-up.
And then you may remember much that will help you.
In public and in private, if you feel surprised or confused, notice your confusion. These feelings are signs that your world model doesn't match reality.
Real incidents make you want to act fast. You feel the need to contact journalists, tweet about the incident, and start telling your friends: a memetic attack will feel the same. If you let them trip you, you burn credibility. Set a 5-minute timer and write out your thoughts before acting. Ask yourself:
Since I haven't sanity-checked the concepts below, please take them with a grain of salt and read the comments where people will likely point out my mistakes.
"What about the German Wiki Attack?": The GWA was found by a team of researchers that we trust, and there was a bunch of public evidence that it had actually happened. I anticipate something similar to this being the trap that is set, but I do not expect it to come from a team of researchers that we trust, and I do not expect there to be more than 1-2 pieces of public evidence.