Can Anthropic or the API figure out which user generated the text?
No. Anthropic confirms in the FAQ that this method cannot do that.
I see no reason to believe that Anthropic is watermarking outputs with user identity, but "this method cannot do that" seems like too strong a claim? If I'm understanding the math correctly, the technique as described can't embed a multi-bit message, but the generalization seems pretty straightforward, if the watermarked text is long enough. A user-ID watermark would be able to coexist with an is-LLM watermark, and the only externally-detectable indication that a second watermark (with identifying information) was present would be that the primary watermark was weaker; and I would not expect external parties without the secret key to be able to detect the weakening.
(Of course, this would be a pointlessly indirect way of tracing outputs to users, since if they were willing to lie about their watermarking procedure they could just as easily lie about their data retention policy.)
If the method is as described (one global key), then "cannot do that" is correct mathematically, since the watermark is a function of the key and preceding tokens. If they are lying and use keys derived per user, it's straightforward to identify the user. For N users, the text required to distinguish them grows like ln N, so for ~4 billion candidate users you'd need ~4x more text to identify the user relative to a single is-AI bit.
But this is externally checkable: if watermarking is as described, then the rate of next-token agreement on identical prefixes should match within an account and across accounts. Higher within-account agreement would suggest keys differ between accounts.
But this is externally checkable: if watermarking is as described, then the rate of next-token agreement on identical prefixes should match within an account and across accounts.
With features like memory and cross chat history, getting identical prefixes across accounts might be difficult. And isn't username part of the context too?
If you added a per-user private key into the random seed, you could indeed check whether a specific user used the same AI to write a block of text.
…but asking which user (if any) used that AI to write the text would be quite expensive. The only way to test that would be to rerun the watermark test for every individual user, and you’d probably have a few false positives in the mix just by sheer numbers. A one-in-a-million fluke should happen about once per million users, after all.
Does this make LLMs much more deterministic? Suppose I ask an LLM to generate a D&D character for me, and do this several hundred times to make a bunch of random NPC adventurers.
An unmodified LLM might make some random decisions and produce a hundred different D&D characters.
A watermarked LLM will watermark itself by reliably producing a half-elf wizard named Erastus, and will produce the same character one hundred times. This will, indeed, let you reliably tell when someone has had an LLM generate their D&D character, as it'll always be a half-elf wizard named Erastus.[1]
This is a relatively minor example, I agree that for most cases it doesn't really matter because the AI randomness is usually not a feature. I do think 'people have a strong default belief that EU mandates around technology will be bad' is really quite reasonable of them in general, though, even if in this particular case the effects are minimal.
I have actually recently had this problem while trying to get LLMs to design and populate a D&D megadungeon for me. I am suspicious that I may have now discovered why.
Yeah, it makes the LLM more predictably skewed. You'd expect less diversity if you were sampling to get the best-of-N. The Synth-ID paper actually literally says this: "For our experiments, we configure SynthID-Text to be single-sequence non-distortionary; this preserves text quality and provides good detectability, while having some reduction to inter-response diversity."
You could quantify how much reduction if you got the hparams / exact implementation from Anthropic, but ofc we don't have this because labs don't show shit. So if you're using an ensemble or multiple calls or a parliament of LLMs you might see worse performance, degree currently unknown.
Hmm, I just tried with ChatGPT 5.6 Luna "Generate a D&D character for me.", and got:
Claude Sonnet 5 seems to like half-elf rangers more, although has slightly better first name diversity:
I don't see sufficient reason to believe this or other common LLM names like Elara Voss are due to watermarking in particular. Isn't it more that LLMs are just generally bad at making random decisions? See for example https://www.lesswrong.com/posts/t9svvNPNmFf5Qa3TA/mysteries-of-mode-collapse and https://maxread.substack.com/p/who-is-elara-voss
I believe it makes it more deterministic, but not much more deterministic. It knows the entropy budget and only consumes a small amount.
I do not believe that it is theoretically necessary for it to be more deterministic at all. If it were deterministic, it would be recording not just the watermark that it is produced by LLM, but also the fact that it was produced at the beginning of the session, which is not the goal and thus wasteful. The purpose is that the watermark can survive the prefix being truncated, so there is a potentially arbitrary context that the watermark should not be sensitive to. It is about correlations between successive calls to the RNG, but it could still be initialized randomly. That initialization is like the unknown context.
Does this make LLMs much more deterministic?
If I understand right, I think it needn't, at least for some prompts? (Which isn't the same as saying it doesn't as-implemented. Also I am not confident I understand right.)
When generating each token, the prng seed used is a hash of the last four (configurable) tokens plus the secret key.
If you're generating something like <user>Write a poem</user><think>User wants me to write a poem</think><assistant>O, codswallop, how I beseech thy shrivelled kidneys</assistant> then the text you can expect to get run through the detector is only what's inside <assistant> and maybe <think>. You won't be asked to check a poem</think><assistant>O. So at a poem</think><assistant>, you can actually just use a random seed, and then you'd get the same amount of diversity in the initial four tokens as you would have without watermarking.
Doesn't help if the prompt has low diversity in the initial four tokens, like "count to four and then write a poem", or if two different contexts with the same last-four tokens otherwise sync up in terms of what they predict.
In theory they could implement it in a way where it makes outputs deterministic like this, but there are easy workarounds to this problem, and the problem would be all-or-nothing in a way that would make it easy to notice. It's only a problem if every bit of available entropy is used for watermarking, so leaving any unused randomness behind will produce a difference early in the response, which will effectively reroll the rest of the response. And if they leave the hidden chain of thought unwatermarked, they don't even have to sacrifice any watermark strength on the non-hidden portion.
Does this make LLMs much more deterministic?
it isn't perceptible by humans (and generation "temperature" already largely controls the level of randomness)
it'll always be a half-elf wizard named Erastus
it's more like, there's an imperceptible bias in word choice in the character's back story: "Erastus's warm, cherished childhood" vs. "Erastus's bright, cherished childhood" vs. "Erastus's rich, cherished childhood."
it isn't perceptible by humans
You will note this is a much weaker claim than "the math is math" and that the output is not even marginally worse.
There are situations where you want to create a text that has a certain vibe and there's a human perceptible difference between "warm" and "rich" childhood.
sure. I was really just trying to create distance between "A watermarked LLM will watermark itself by reliably producing a half-elf wizard named Erastus, and will produce the same character one hundred times" and what text watermarking actually looks like. in a real system they have lots of entropy to play with. it's part of the worry with steganographic communications by models between each other.
I oppose because I think it's inevitable that these things get integrated and outputs end up being collaborative, and I don't trust the social stigma that will develop on this. I'd rather outputs be judged on their output quality, and I don't like the incentive becoming all weird wrt how to achieve it.
There is basically no handle here that won’t have the wrong vibes in the same way, and give the impression that you did extra work to change something, and therefore made things worse. I think watermark is a relatively good name and would prefer not to go on a euphemism treadmill.
I would suggest "fingerprint".
My response is that this is going to be annoying, and also doing it is clear consciousness of guilt and far worse than the initial AI use
I'm remembering how go players disempower themselves to AI:
we learned from the many examples of cheating and player confessions that idle curiosity and laziness were the dominant reasons for AI use in our school. Our students would often set out to play a normal game of Go, but would get stuck on a particularly difficult or annoying move; eventually, their curious eyes would drift to their second monitor — where they usually had their AI software running anyway — and they would check the answer as one would sheepishly side-eye the solution to an interesting puzzle or homework problem. ... This perspective of AI use to me explains why camera controls proved so effective against online cheating. Since AI use is usually an act of self-debasement and disempowerment – a subjection of oneself to ambient incentive gradients – it fundamentally contradicts the aesthetics of resourcefully overcoming a minor obstacle.
Isn't the cost here that you need to keep track of the seed(s) used in your pseudorandom number generator and store them as a secret key? It does seem like a fairly trivial cost, though.
And does this method work if you don't have the entire AI output from the beginning? For example, you ask the AI to write ten essays in a row, and then turn in only the last one to check if it has the watermark. If your PRNG is good, can you easily tell if a sequence "eventually" gets produced by a seed? Or do token limits make this a non-issue?
DRM … makes your product worse. At best it eats resources, and it can actively prevent you from using the product, and sometimes it can mess up your machine. All us gamers know the pain well.
Have you seen Irdeto's marketing for Denuvo? They also claim that the performance change, when they're not saying it's literally zero fps, that it's within run-to-run variation, that properly implemented, it's perfectly imperceptible for players, that the only reason to have any concern is that you want to pirate the games they protect, etc. If one were to accept their assertions with the same level of credulity that you do Anthropic's, you'd recognize that they believe absurdities.
Scott Aaronson, while working at OpenAI, largely solved AI text watermarking together with Hendrik Kirchner.
Here is how his solution works, or see Tenobrus’s version.
If you want to dig deeper, here is a full paper. The method has very nice properties:
The European Union Code of Practice, signed by the major Western AI labs, requires future AI models to use such watermarks.
Google implemented this, including for Gemini 3.7 Flash, and they have been rolling out this feature since 2024. Google has done, for over two years, the exact thing Anthropic is now doing, except with a public detector, and Google confirmed in a test (n = 20 million) that there is no difference in user feedback.
Anthropic quietly announced a week ago they were rolling out watermarking to comply with the EU Code of Practice. Since they don’t want to have to differentiate traffic sources, the marginal cost is zero, and watermarking is pro-social, this will apply to everyone. They then offered an FAQ of how it works.
It is possible that, once they have the ability to differentiate for other reasons, they will use it here as well, if we decide universal watermarking is bad. I think it is good.
My initial read was that this was a quiet positive story of a good thing, showing that if something good worked with zero downsides or costs then maybe we would do it, so this was my full initial coverage:
This is distinct from studying removal in order to account for or defend against it, which is obviously fine. You’re only asking about being the Baddies if you’re actually removing them in practice.
Table of Contents
This Is Fine
No. Not so much. A lot of people responded by getting Big Mad.
So here we are.
The entire practical effect is: There will be an API that will tell you if a given piece of writing comes from Claude. That’s it. And yet.
The rest of this post is about exploring why people are Big Mad about this, in large part as a worked example of how people get worked up over approximately nothing.
My conclusion is that a bunch of different factors are coming together.
Anthropic Derangement Syndrome
This is the main reason. Let’s not pretend otherwise.
You don’t see people getting Big Mad at Google over this. You don’t see them getting Big Mad at OpenAI, even though that’s where this was invented and they have committed to doing this going forward. And so on.
It is unfortunate that Anthropic was the first to announce they were implementing watermarking to comply with the EU Code of Practice.
Because this is now associated with Anthropic, all sorts of bad vibes try to attach themselves. Certain types of people look for reasons to be upset.
Anthropic can’t win. If they were initially louder about it, they would tie watermarking to Anthropic. Because they started out insufficiently loud, due to this not actually being a big deal, people get mad about that instead, and then still tie it to them, despite Google having implemented it and shipped a public detector over two years ago.
AIs are not widely considered people, but credit should go where credit is due.
Are some of the other complaints about watermarking legitimate, understandable or born of genuine misunderstandings? Sure. But a lot is that people think Anthropic vibes are bad, and thus look for reasons to be upset, and on principle refuse to believe the explanation I put up top, and assume something sinister must be going on.
Some people’s paranoia and derangement is directed at abstract notions of ‘openness’ rather than Anthropic in particular, which amounts to the same thing in context. This is an example of fetishizing that this is not ‘open’ therefore must be sinister, even though in this case actually it is mathematical and anyone could verify it. By asking closed model Grok, of course, because Elon Musk has vaguely open vibes despite keeping all its competitive models closed.
People Don’t Understand LLM Outputs Are Already Random
If AI outputs were deterministic, it would be impossible to encode a watermark without making them at least marginally worse.
Many people intuitively think that the AI outputs are not random. That what they get is the One True Output, even though you can regenerate the output and it will reliably be somewhat different. Thus, if you’re not getting the ‘real’ or original output, that means your output must have gotten worse.
When people hear watermark, they think it will make the outputs worse, because to leave a mark you have to make different choices. You have to change something.
Similarly, they would assume that if you are optimizing for an additional thing, it is going to cost more.
This is a good intuition.
It turns out to be wrong here, because you have enough randomness to play with that you can get the same effective distribution, and still encode the watermark.
And no, everyone does not know how this works, almost no one reads Scott Aaronson in detail because almost no one reads, very few things are actually common knowledge. If you follow discorse expecting people to know basic technical facts you are going to keep being deeply confused.
People Don’t Trust The Method To Be Costless
I believe the skepticism here is greatly enhanced by Anthropic Derangement Syndrome, and by general distrust of Anthropic, a sense that ‘they’re up to something.’
How much of the skepticism is due to skepticism of the method?
A quick survey suggests that this is the majority of the concern.
This thread is an example of someone finding it difficult to accept that there is effectively no impact on outputs, because the distribution does not change.
I do sympathize. This is a magician’s trick, a math proof, that works. There is something highly counterintuitive about ‘you can mark it while having no impact’ and every fiber in people’s bodies wants to say no, until something clicks and they realize that the math is math and actually yes it works.
I will think actively better of you if you ‘take the L’ like this in public, even if I never saw the original L. There is an obvious game theoretic issue with that if everyone predicted everyone would react that way, but they don’t, so this play is safe and wise.
This is an example of someone trying really, really hard to say that this method technically ‘has tradeoffs’ to imply it is not costless, because there is a strong drive to not want it to be costless in order to be mad about it, when obviously in practice it is costless, indeed Google ran extensive experiments to prove it is costless.
This is an example of flat out ‘nope, I don’t believe it’ on principle, despite the math being very clear. And this response is an example of hallucinating a loss in quality, that is claimed to be observed, based on logic of ‘well I think hard about my choices’ and refusing to understand the choice is random either way.
This is a (much more egregious) example of someone pretending not to understand what random means, reading ‘there are two possible continuations’ as a claim that the two continuations mean the same thing.
People Are Suspicious Of Any Alteration On Principle
From the above survey:
There is a lot of this kind of attitude, things like:
Except of course, whether the model be open or closed, it is the product of thousands or millions of decisions, many of which are not about what you the customer wanted. All such products are ‘altered’ constantly. Often the change will not ‘serve your needs,’ either yours in particular or those of users or customers in general. There are many other problems to solve as well, including legal requirements and harm prevention.
When this is invisible, people do not care. When one in particular becomes salient, people get angry.
The parallel to DRM is also poisoning the well here. DRM sucks, in all its implementations, because it makes your product worse. At best it eats resources, and it can actively prevent you from using the product, and sometimes it can mess up your machine. All us gamers know the pain well. You sometimes have to do some of it, because copyright does not enforce itself even if you in particular would still honor it.
This is not like that. But it slightly vibes with it, and that can be enough.
Maybe It’s Partly The Word Watermark
I mean, I guess? I presume ‘secret’ would be ten times worse for the same reasons.
There is basically no handle here that won’t have the wrong vibes in the same way, and give the impression that you did extra work to change something, and therefore made things worse. I think watermark is a relatively good name and would prefer not to go on a euphemism treadmill.
A Lot Of People Don’t Want To Get Caught
Of course, you can’t come out and put it like that. Well, some people can, but a lot more of them choose not to. Often this will not even be fully conscious.
Don’t overthink the situation. A lot of people value the ability to pass off AI writing as their own, or they want people to be unable to prove it.
And no, you won’t be able to switch to ChatGPT, they will have marks too.
There is also a bunch of ‘not that you would, but you could.’
There Are Some Times You Prefer Not To Be Recognized
Can Anthropic or the API figure out which user generated the text?
No. Anthropic confirms in the FAQ that this method cannot do that.
There Are Some Good Reasons To Be Concerned
With any change that impacts social dynamics, even when the objections are primarily wrong, and the thing seems clearly positive, there are still going to be some downsides and concerns.
Here are the ones that seem at least somewhat legitimate to me.
Cheat Cheat Cheat Cheat Cheat
This is my top real concern, which is that the worst people will remove the watermark.
Removing the watermark is non-trivial, but given you have an answer key to train with and check in each case, one can doubtless build AI tools that scramble small choices in ways that degrade or erase the watermark. You can verify that it worked.
They could also use a local or other model that does not carry a watermark, or not one that anyone is likely to check.
This potentially puts you in a worse position, since they can then use the false negative as a defense. When you have a test that is good enough that you trust it by default, but that is possible to fake when it counts, that can be pretty bad.
My response is that this is going to be annoying, and also doing it is clear consciousness of guilt and far worse than the initial AI use, and that other methods of detection will still work. Pangram will still mark it as AI, at least if you do it the way I’m imagining, and won’t be confused about which AI the original came from (the Pangram algorithm can differentiate different AIs, but it doesn’t tell the user). It would also still read to a human as AI.
Thus this would be an issue if for example a college student used it, and the university couldn’t act without the proof from the watermark, and that issue could snowball, but now the student needs to put in more effort, and has clear mens rea, and in practice I expect this to be mostly not something that is done.
The real world test is another reason not to worry. Everyone has access to Pangram. So in theory, anyone could iteratively check the Pangram result until it comes back as human. But we have observed how people react in the wild, and approximately no one does. People just… get caught.
In general, if you raise the effort level of submitting AI work, that mostly does do the job. If you remove the watermark by rewriting the whole thing in your own words, of course, that is fine and counts as Mission Fucking Accomplished.
The Writing In The Middle and Error Rates
When writing uses AI to some extent, but is still largely written by a human, what happens with the watermark? If you use some of the AI turns of phrase, will people classify your work as AI, even if it is largely your own? Will there be zero tolerance policies and unfortunate cases? This is all probabilistic, so what happens when the answer comes back wrong?
The answer is that the watermark measures AI processing. So translations and file conversions might trigger the watermark, which is unfortunate, but one can be aware of that. So could proofreading, if you let Claude automatically make the related edits, but if you use it to find errors and correct them yourself, it won’t.
Thus I am confident that strong watermark positives will consistently be true positives in terms of AI writing or processing of the words. The amount of watermark signature on my posts will not be zero, since I am quoting others who sometimes will have used Claude, or sometimes directly quoting Claude. The direct quotes are clearly marked, but the watermark won’t know the difference. There will be some degree of paranoia when using any words from an AI system, in places where such use is clearly good.
So yeah, a little of that will happen, but I expect people to rapidly get used to this, and there should be a high presumption that giving people more info is good and it is up to them how to react to it. If you are paranoid and among people who are Big Mad about even a sliver of AI use, and also that might actually use the API, you can check your own output first via the API. Or you can honor the preferences of those folks, and not use AI even in some of the ways that you and I would agree are good.
As usual, in terms of actual mistakes and false positives, people have vastly lower tolerance for errors and potential errors by automated systems, than they do for humans. The error rate is going to be very, very low. If a human is trying to decide if you used AI, they’re going to have a substantial error rate, far higher than the watermark. And AI use is one place where demands to ‘prove’ things go too far, especially in academia, often letting people often ‘get away with’ things that everybody knows they did. This is not criminal law, if no one is going to jail you should not need the same super high level of confidence.
Millions For Defense But Not One Cent For Tribute
The last concern is not about the watermark itself, but about the mandate and its origin in the EU Code of Practice.
Anthropic implemented watermarking worldwide due to an EU law, since it is a lot easier to do it everywhere than only for the EU. So that means that the EU is sapping and impurifying our precious AI output tokens, against our will. We must fight back. Who knows what else they might target next?
The object level response is that no, they are not doing that. They are technically changing the outputs, the same way that a butterfly flaps its wings and changes the weather, but not in any systematic or directional way. As discussed above, the outputs are not degraded. This Is Fine.
The real concern is the principle, and what might come next. Yes, this is fine for now, and forcing Apple to use USB-C was fine for now, but the fines they impose on our tech companies are basically modern piracy and who knows what comes next.
To which I say, they are a huge market, and yes they get some say, and have always gotten some say, and the limiting factor is America pushes back or in extremis we geofence, take our ball and go home, as indeed has already happened with some AI services, in the EU and elsewhere, over similar issues.
Another thing that might come next, in theory, is that there is no size minimum on requiring watermarks. So in theory they could come after any AI without them, including tiny open models. If they made a sufficiently large fuss about this it would be bad. I do not expect this, but they’ve done stupider.
This is the dance. I am worried about the EU or others imposing a censorship regime in this way, forcing our tech companies to play along. That could plausibly extend to ideological requirements on AI outputs. If they or others did try that, I believe our frontier labs would not apply such changes globally, and indeed would face severe backlash if they tried. Instead, I predict the labs would at least threaten to geofence, or use aggressive classifiers or similar tech on EU queries.
Yes, we do have to keep an eye out for Brussels overreaching. But this? This Is Fine.