SereyJones has been concerned for a while now that the avalanche of AI generated text and complete books will dumb down and remove the human DNS from prose. We have seen ads offering to "publish your book with AI" and from "idea to published book". Here at SereyJones, we are very much opposed to this method of generating this AI slop. It appears that three AI large language models (LLMs) also want to distinguish human hard work versus lazy AI button pushing. I have collected below what the top three AI companies are doing to make it possible to detect AI generated text.

Artificial intelligence platforms like Claude, ChatGPT, and Gemini embed invisible statistical "fingerprints" into generated text to help identify whether content was written by a human or a machine.

The Challenge with Text
Unlike digital photos or videos—which can easily hide metadata inside their image files—plain text consists only of standard characters. Copying and pasting text strips away file metadata, while adding hidden characters can ruin grammar or distort meaning. To solve this, AI developers slightly tweak how the AI chooses its words.

How Invisible Watermarking Works
Imagine an author writing a story while secretly following the digits of Pi:
- Every time the author needs to choose between valid synonyms (like "happy," "joyful," "cheerful," or "delighted"), they check the next digit of Pi.

- If the digit is even, they pick a word from an approved "Green List".

- If the digit is odd, they pick a word from a "Red List".

- If the preferred choice sounds awkward, the author overrides the rule to keep the sentence natural.

To an ordinary reader, the final story flows completely naturally. However, an auditor who knows the secret Pi rule can analyze the word choices. Finding that 85% of choices align with Pi's digits—where a human would only match 50% by pure chance—proves the text was written using the secret rule.

How AI Companies Implement It

- Early Research (2022–2023): Researchers at the University of Maryland proved that AI models could secretly favor certain word patterns without degrading text quality.

- Google DeepMind (Gemini): Developed SynthID, a watermarking system that maintains high text quality and has been open-sourced for developers worldwide.

- OpenAI (ChatGPT): Tested secret key patterns that achieve over 99.9% detection accuracy on passages longer than 200 words.

- Anthropic (Claude): Researched how well watermarks survive when text is edited, rephrased, or translated into other languages.

0 Comments

Submit a Comment

Your email address will not be published. Required fields are marked *

I accept the Privacy Policy