How Claude's text watermarking works
Aug 18, 2026
There has been a lot of discussion about Anthropic's proposed 'watermarking' of AI-generated content. This article explains how it works, which in turn informs our intuitions about why it's so objectionable. Basically, what the system does is to subtly alter word selection when generating text, creating a difference in the probability of what it outputs as compared to what it would ideally output. This difference is essentially a 'signature' attributing the content to a certain AI. Anthropic will later offer an API which allows people to test for thaat signature. The problem is, if you use AI, the AI is messing with your text to embed hidden messages in that text, which is (to say the least) problematic. John Gruber calls it patently offensive. Tim Moon and others calls it a digital scarlet letter.
Today: Total: [] [Share]

