Blog
Why Claude’s new AI watermarks can’t replace AI writing detectors
The reason we need AI writing detectors alongside AI watermarks is that they do very different jobs, and while traditional AI text detecting services aren’t perfect, neither is AI watermarking.
Let’s start with Claude AI watermarks, which I’ve covered repeatedly over the past couple of weeks. Once it’s fully up to speed (the newest Claude models will support watermarking first, while older models will be retrofitted later), Claude will use a version of Google’s SynthAI watermarking technology to “nudge” the sampling of certain words it generates, typically “low stakes” words where its probability engine gives it a range of viable options.
Welcome to another edition of Prompt Mode, your weekly AI newsletter.
I’m your host, Ben Patterson. Each week on Prompt Mode, I’ll be serving up analysis of the AI trends that matter to everyday users like you and me. Stay tuned for practical AI tips, hands-on experiences with the latest AI tools, and–you guessed it–prompts to help you get the most out of your AI assistants.
Thanks for reading, and if you like what you see, just sign up right here.
Strung together in a long enough passage (roughly 150 words or more, or at least 200 tokens worth), these “nudges” form an invisible pattern that’s detectable with the right tools, which Anthropic says are coming soon (and will presumably be free). The presence of a Claude AI watermark definitively tells you that Claude has “processed” the text in some way, either generating the original words or substantially editing them after the fact. (“Light” Claude copyediting won’t trigger watermarks, according to Anthropic.)
But while the presence of Claude AI watermarks will tell you for sure that Claude has touched a piece of text, their absence won’t prove that it didn’t. While Claude watermarks will survive a straight cut-and-paste, they can be disrupted or erased if the words are heavily edited, either by a human or another, non-watermarking AI model. Also, Claude AI watermarks only appear in Claude-processed text, although it’s possible other AI providers will employ similar and perhaps compatible AI watermarking technology.
Whereas Claude watermarks are akin to tangible strands of DNA, traditional AI detectors are like polygraph tests. A polygraph can’t prove whether someone is lying (which is why they can’t be used as evidence in criminal trials), but they can detect subtle changes in respiration, blood pressure spikes, increased sweat, muscle tension, and other physiological signals often associated with lying.
Same goes with AI writing detectors. While the can’t tell you for sure whether a passage of text is AI generated, they can spot writing patterns suggestive of AI generation, such as strings of words that are popular with AI models, semantic “looping” (where an AI model starts repeating itself), sentences that lack variety in terms of length and rhythm, and of course, the classic AI “tells”: too many em-dashes, phrases like “it’s not this, it’s that,” lots of formal transitions (“furthermore,” “in conclusion”), and a plethora of hedging phrases (“it’s clear that,” “it’s worth noting”).