[$500 Bounty] I'm offering a bounty of $500 for someone with red-teaming skills to build attack LLM pipelines for large-scale online deanonymization. — LessWrong
In the above experiment, researchers from Anthropic and ETH Zurich were able to build attack pipelines to essentially deanonymize Reddit users using a combination of text patterns (aka a sort of "writer's DNA") and contextual clues (i.e. 35 years old, works in tech, lives in San Francisco, etc.)
Obviously public and non-public models will continue to improve at these capabilities, but I am looking to see if this technique can be replicated using existingmodels to de-anonymize any of my accounts with my public blog posts.
I will offer a $500 bounty to someone who is capable of building a similar pipeline, and will offer $1000 total if they can uncover any of my anonymous Reddit accounts.
For someone to participate, I'd need:
Proof of strong red-teaming skills and/or large-scale LLM pipeline building with a strong understanding of Stylography and related concepts
Assurance that they will keep any recovered findings private
Presumably, you want to have a copy of the pipeline, or have it openly published. Which means you’re commissioning a general use weapon. I’m sure such tools will eventually become widely available, but I don’t see any benefit to the world to hastening the process.
We also have no way of knowing whether the Reddit accounts you’re deanonymizing are actually yours.
Not necessarily having it openly published; I would be more than happy with someone who red-teams in good faith building this and reporting to me the results they found.
And I might not have been super clear in the OP. The goal isn't to de-anonymize reddit posts, but to 'crawl' for reddit posts that contain similar text patterns and contextual clues that are present in my (public-facing) writings.
How can someone know whether the public facing writings you give them are yours? The only identity you have here is two posts and two comments, no proof of identity attached.
Also, this process latently deanonymizes a bunch of strangers too, but for someone pointing the pipeline in their direction. Can you trust the bounty hunter to not want to use this weapon for themselves, or share it with others?
The results of that paper only gives a probability. Would you accept that from your bounty hunter? Or would you want them to give you the pipeline?
What do you need this for? The paper proves it can be done. You know you are already going to be deanonymized with everyone else sooner or later. This won’t give you any new evidence about whether e.g. you should tighten up your online presence.
I will obviously prove my identity later on and give text from elsewhere. I obviously do not intend to use my two posts and two comments from LessWrong.
I think someone who would build this tool for nefarious purposes is not going to suddenly jump from "not doing nefarious things" to "doing nefarious things" due to a $500 bounty, that seems unrealistic.
I already said that I would offer the initial bounty to someone who replicated the results of the paper but clarified it would be doubled if it surfaced my actual posts (which I would verify). That would be a binary outcome.
I would like it to know if my public-facing posts are sufficient to detect posts elsewhere (assuming they too are sufficient). It would in fact inform me that I could alter my stylography on new forums in the future to avoid them being attributed to texts elsewhere.
Sorry, I pressed a lot there. I’m just worried about what this does to the world. Legitimizing this sort of thing corrodes the commons, and the world is a slightly better place if the commons corrode a little slower.
The whole exercise also relies a lot on trust, which can easily be broken. It could be something as simple as, ”the bounty hunter decides this would make a good portfolio piece and puts it on GitHub.” Or they share it, “look at this cool thing I made,” thinking they’re being a legitimate privacy researcher who got paid real money by a conscientious client, and it spreads from there to people who aren’t. Or even, “did you know you can build a deanonymization tool in your spare time and it only cost this much? No, I’m not going to share mine,” but now everyone knows it can be done.
You could probably still get what you want by running all your output through a stylometric randomizer now. Given the paper, you will want to have been doing so anyways by the time this eventually becomes a common weapon.
Large-scale online deanonymization with LLMs
In the above experiment, researchers from Anthropic and ETH Zurich were able to build attack pipelines to essentially deanonymize Reddit users using a combination of text patterns (aka a sort of "writer's DNA") and contextual clues (i.e. 35 years old, works in tech, lives in San Francisco, etc.)
Obviously public and non-public models will continue to improve at these capabilities, but I am looking to see if this technique can be replicated using existing models to de-anonymize any of my accounts with my public blog posts.
I will offer a $500 bounty to someone who is capable of building a similar pipeline, and will offer $1000 total if they can uncover any of my anonymous Reddit accounts.
For someone to participate, I'd need:
Thanks