
There has been a growing tug of war between publishers and
authors over AI. Until recently, I thought we were slowly beginning to figure out some reasonable ground rules.
If you ask an AI to write your book and put your name on it, that's one thing.
If you write the book and use AI to check spelling, grammar, citations, chronology, repeated paragraphs or whether you've gotten somebody's name wrong, that should be something very different. The
line between those two things can get blurry, but at least we’re beginning to acknowledge that there is a line.
Anthropic may have just made the whole thing worse -- maybe on
purpose?
Here’s where the whole kerfuffle begins. Claude is going to watermark all the text it generates, embedding an invisible statistical signal into the words
themselves.
The marking is not limited to Europe. Anthropic said it applies everywhere Claude is offered, worldwide. No opt-out.
advertisement
advertisement
Anthropic says it is doing this to comply with
the EU AI Act, and on the surface this sounds perfectly sensible, right? We should absolutely know when text was created by AI. Except that's not quite what the new watermark does.
In an
online video, Casey Fiesler, an AI ethics professor at the University of Colorado Boulder, provided a very good explanation of how this probably works, and got to the problem pretty quickly. “A
positive probably means that it was touched by Claude, if not written entirely, then edited or paraphrased,” she said. But a negative doesn't prove the opposite. As Fiesler puts it, “A
negative detection doesn't actually tell you much about whether Claude touched that work,” because enough rewriting can make the signal disappear. Confused yet?
So a publisher may
be able to learn that Claude was involved. What the watermark can't tell them is what Claude actually did. And that's the question that matters.
I was thinking about this while listening to a
fascinating conversation between author Steven Johnson and Nicholas Thompson, CEO of The Atlantic, on Thompson's podcast “The Most Interesting Thing in AI.” Johnson has written 14
books and is now editorial director at Google Labs, where he helped develop NotebookLM.
Johnson said that he has been using software to help with his writing since college. Today, he uses
NotebookLM while writing a book about the California Gold Rush, with roughly 200 sources and millions of words of research available to query. Instead of spending an hour hunting for something he
remembers reading, he can find the relevant source passage in seconds.
Is that AI writing the book? I don't think so. Neither does Johnson.
When Thompson pushed him about where the
line should be, Johnson described his own boundary pretty clearly: “What I want is I want it to help organize all the information and present it to me in exactly the way that I need to see it at
any given time,” he said. “I’m happy doing the judgment part of it.”
Thompson offered an even better example from his own experience. He recently wrote a book and used
AI “all the time,” including to check chronology in complicated narratives. But Thompson established a strict rule for himself: “I didn't want a single word in the book to have been
written by an AI model.” He said If the AI found a problem, he rewrote it himself. Does that escape the Claude watermark? Hmm… maybe. Who knows.
But that strikes me as a
pretty defensible definition of authorship. But here's where the watermark question gets complicated. Imagine an author writes every word of a book and then uses Claude as an editor. Maybe Claude
fixes some grammar, tightens a few sentences or suggests a better word here and there. None of the research, reporting, ideas or argument come from Claude. But some of the words now do.
When a
publisher finds the Claude watermark, suddenly we have a problem that didn't exist before. The author says, “I wrote the book.” An AI detector might examine the prose and say it appears
human-written, while Anthropic's watermark says Claude was involved. All three statements can be completely true.
For a publisher, “Claude watermark detected” is going to sound
considerably more definitive than an author's explanation of how the software was used. A watermark reads like forensic evidence.
That's especially troubling because publishers themselves have
started drawing these distinctions. Elsevier, for example, says authors don't need to disclose AI use for basic grammar, spelling and punctuation checks, while substantive generative uses do require
disclosure. That's a reasonable policy because it focuses on what the machine did, not simply whether a machine was involved.
Claude's watermark threatens to bring those two things back
together.
There is also a strange incentive problem. Someone who generates an entire article or book with AI and wants to conceal that fact has every reason to remove the evidence. They can
rewrite the text, heavily edit it or simply use a system without the same watermark. And, as Fiesler puts it, “significantly rewriting probably destroys” the watermark.
Legitimate
authors have no reason to do any of that. They wrote the book and used Claude as an editor, accepting some suggestions and rejecting others. They're not trying to hide anything, so some
Claude-generated language, and potentially its watermark, stays right where it is.
The cheater scrubs the fingerprints. The author leaves them behind.
Johnson offers a much better way
of thinking about this. Imagine, he says, having “a world-class researcher and a world-class tutor and a world-class editor” sitting at the table with you. “Would you go to your
editor and say, ‘Write this paper for me?’ No, you would never do that.” His rule of thumb is wonderfully simple: “If you would feel sketchy about asking a human in that
capacity to do that work for you, that's probably something you shouldn't be doing it with AI.”
Johnson's distinction makes sense. The important question isn't whether AI was somewhere
in the room. It's what we asked it to do. Did it check the spelling or write the paragraph? Did it locate the source or invent the source? Did it identify a weak argument or create the argument?
Those are the distinctions publishers should care about. A watermark doesn't tell us any of them.
And an invisible mark that says only “Claude was here” may create a different
problem. It gives us certainty about the least interesting question, while leaving the important one unanswered: Who actually wrote the book?
And, just to be darkly conspiratorial, what if
Anthropic’s actual mission is to continue to make the distinction between “AI-assisted” and “AI-drafted” so impossible to define, that writers, publishers, and AI
detectors are all painted into a corner? From a purely business perspective, that would be the ideal outcome, wouldn’t it?
And consider how the Anthropic
watermark arrived -- not with an announcement, but as a quietly updated support page that users found on their own. If the goal was transparency, transparency about the watermark would have been a
reasonable place to start. And if you're a company that planning to sell provenance as a service, a world where nobody can tell "AI-assisted" from "AI-drafted" without your tool would be the
ideal outcome, wouldn’t it?