AI watermarking text

It’s happened sooner than I expected:

This has blown up in the blogosphere.

Anthropic: How Claude marks AI-generated content | Anthropic Help Center

John Gruber: Daring Fireball: Anthropic’s ‘Watermark’ Text Adulteration in Claude Is a Perversion of Writing

How it works: Anthropic’s weak watermarks appease a weak law - by James Padolsey

Gruber again: Daring Fireball: 'Anthropic’s Weak Watermarks Appease a Weak Law'

Dan Moren: LLMs aren’t writing – Six Colors

More from Gruber: Daring Fireball: Follow-Up Thoughts on Watermarking Schemes for AI-Generated Text

I don’t have a strong opinion about this in either direction. I find most AI-generated text cloying and award mental demerits for using it unless English isn’t the author’s first language. But I also find poor human-written text annoying, and award similar mental demerits for it.

I don’t quite get Gruber’s take. Using AI for your writing is the problem, not the idea that AI might watermark it.

2 Likes

Yeah, there’s a lot to unpack on all sides, which is in part why I haven’t really formed an opinion yet. I don’t quite get why watermarking is definitively bad for Gruber (does it make the text worse or just different, and if no person is behind the specifics, what does it matter?), nor do I quite agree with Dan Moren that LLM-generated text isn’t writing. It feels like a bit of a Turing Test—if you can’t tell it was AI-generated, why does it matter?

1 Like

I’ve also been thinking Gruber’s take is a little extreme.

We can agree that human writers have styles? Writers don’t always pick the absolutely perfect word every time. For example, I don’t use shibboleth even when it would be the best word.

This is why stylometry works.

So why shouldn’t an AI also have a writing style, where it tends to use some words more than others? I don’t agree with Gruber that it is degrading the output for it to do this.

1 Like

It is odd that LLM produced text is only able to be determined to be LLM by another LLM. It is not detectable by humans alone.

I do agree that his thought on “degrading” is extreme – AI “writing” is already poor, so making it .01% more poor isn’t a big deal. I care about my words so I would never let AI change them.

But there are three key issues that are critical:

  1. This “ai detection” doesn’t detect that AI text is AI-written, only that it’s been AI-processed. That means proofreading, or quoting a line of text from another source (or AI) could raise an “AI” flag on your 99% human-written text.
  2. There’s no answer as to what we do about this. So what if text is AI-processed? Do we not use it? Call it cheating and refuse to publish the novel that was proofread by AI? A tool that tells us text is 90% AI-generated might be useful. A simple “AI” flag (on or off) tells us nothing about what % of AI was used on the text.
  3. This watermarking is a way to appease AI-haters and make them think they’ve accomplished something, when it does nothing. Anyone who really wants to deceive with AI will use a non-watermarking AI or run it through Declaude (a filter that makes AI text less AI-like) and those are the people generating AI slop that we want to stop. Watermarking won’t help with that.
1 Like

That’s not quite true. It’s only detectable by the original LLM. Other LLMs use a different watermarking scheme, and since these all use a proprietary secret key to do the watermarking, they are not interchangeable and cannot be detected by others.

Which means, like Gruber says, to detect the watermark you’ll have to run suspect text through dozens of detectors from all the watermarking AI makers, to see if you get any hits!

Based on this, it is clear that the real goal is for an LLM provider to detect when their output is distributed to third parties, perhaps in violation of some term of service.

This kind of watermarking isn’t really useful for any other purpose. But maybe it’s enough to make EU lawyers go away, at least for a few minutes.

I saw one article mention that watermarking might be useful for AI companies to exclude AI-generated text when training. This would be helpful in not degrading AI training by having it absorb more AI text (similar to photocopying a photocopy).

But I’m not sure how well that would work if we end up with dozens of watermarking schemes.

And how will text degrade if people use more than one AI on text? (I currently sometimes use several AI models to proofread my text.) If they each modify the output slightly to add their own watermark, could that degrade the text significantly?

What is interesting about this is that computer-generated text is not protected by copyright, meaning that while distribution might violate something in a TOS it’s definitely not something where copyright applies.

My understanding is that publisher’s paranoia about AI-generated text is in part because of the lack of copyright protection.

Dave

It’s also because if publishers can’t be sure of the provenance of text (or other intellectual property) then they could be publishing plagiarized works, and possibly violating copyright.

It’s why most contracts have a clause indemnifying the publisher if a work violates copyright (which doesn’t mean they won’t be sued or dragged into court etc.).

1 Like

Gruber? Extreme?

At least you know in no uncertain terms where he stands. Which is refreshing.

1 Like

My first thought on the matter was to wonder about unintended consequences.

For example, although well-intended, GDPR and the ePrivacy Directive have left the Internet littered with cookie consent forms. Have these forms actually achieved their intended purpose, or have they simply made the Internet more annoying for most people?

Will watermarking improve traceability, will it increase A.I. sloppiness, or will it be yet another “improvement” whose impact is anything but? Time will tell.

1 Like

His take about many things (but not all) seems a little extreme though.

you can remove AI company invisible text watermarks at www.demarkify.com - and also perform other tonal text changes like removing AI lexical tells

New Scientist has an article describing EU efforts to identify AI generated text:

It seems my ASCII character suggestion hasn’t been considered.