Now that the EU mandated watermarking, the point is that services (or browser extension developers) can add their own detectors to make AI-generated text obvious. It won't fix AI in print, but most of the problem is online anyway.
These things are trivial to remove though. And the whole point of it is they will also make the _detector_ available so you can then also check if you successfully removed it.
An approach like C2PA is the only realistic path forward. If the trajectory we're on continues, it's probably safe to assume nearly all content will be AI generated. We need realistic ways to prove content is human generated, and without true authentication (someone willing to corroborate they created the content, and they can certify it), the whole endeavor is pointless. While private human-verifiable content will still cease to exist, at least in this way we can avoid moving into an information dark-age
True, but what you can do is a one-sided guarantee. If it bears the mark, it is likely generated (or someone deliberately made it look generated).
Thus, if a news article, research article, book, student paper submission, blog post, HN comment, etc, bears the mark, it could be automatically flagged as such.
It helps detect low effort slop.
---
Caveat. If you write your own creative work and send it to Claude for "cleaning up grammar", it might insert the watermark.
The problem with pretending is that people who k ow what they’re doing get away with it while people who don’t (and don’t even use ai) get unfairly accused of using it.
There just isn’t enough information in plain text to do this and we should stop pretending there is.
If we need to verify something isn’t made with ai then we need other ways of doing so - eg looking at a document edit history, doing it as an exam, oral defense.
There are options! But pretending you can tell if text is ai will only catch out people who make no effort to hide it and will inevitably have false positives.
It seems like it would be so low effort to bypass, especially when you can just train a system (maybe even another LLM) using the watermarker validation from Anthropic themselves.
Might help with students and low-technical people passing off work as their own, but any industrial scale slop-generator should be able to bypass it trivially.
Not might, will. Whether enough text is present or not to go over the detection threshold is in doubt. But the "score" will never be zero, even for human written text.
It’s just not a reasonable ask.