Skip to main content

OpenAI details new text watermarking system for ChatGPT, Codex, and the API

OpenAI today announced how it will comply with the EU AI Act, which requires providers of generative AI systems to mark AI-generated text in a way that makes it identifiable. Here are the details.

A bit of context

A few months ago, Anthropic sparked controversy when it announced that, to comply with the EU AI Act, Claude would start watermarking text by slightly adjusting its word choice.

In a nutshell, every so often, Claude chooses between words that would work equally well based on a hidden pattern, which Anthropic says has no practical impact on the meaning, quality, or readability of its responses.

Anthropic also developed a detector that can look for that pattern in a piece of text, and can tell whether it likely came from Claude.

The detector can’t be used to determine whether text came from OpenAI or any other competitor, since Claude’s watermark is generated using a secret key specific to Anthropic’s system.

At the height of the controversy, Anthropic stressed that Claude would not be alone, noting that several other major AI providers were also implementing marking systems to comply with the EU’s requirements.

Today, OpenAI announced its own solution, which relies on a similar approach, albeit with a much more limited rollout and scope.

OpenAI’s text watermarking system

According to OpenAI, “API customers globally will be able to opt in to text watermarking for select models” starting today, while the company “will add an invisible watermark to eligible ChatGPT and Codex text output in the European Union” in the coming weeks.

This means that in the API, text watermarking will remain off by default for now, while in ChatGPT and Codex, it will initially be limited to eligible users only in the European Union.

OpenAI also opened applications for access to its text watermark detector, which will initially be limited to approved researchers and expert organizations as the company works to evaluate and improve the technology.

And speaking of accuracy, OpenAI is very clear that the detection technology still has significant limitations, as it can produce both false negatives and false positives, particularly with shorter texts or content in specific domains, “such as mathematics, where there is less flexibility in word choice.”

Here’s OpenAI on textGrain, which is its text watermarking technology:

Our text watermarking technology, textGrain, adds an invisible statistical signal to the model’s word choices. Our detector looks for that signal to assess whether a passage contains an OpenAI watermark. More details about how textGrain works can be found in our technical report⁠, which will be updated with additional details in the coming weeks. We also plan to make the technology available in open source so that others can build on it.

The company also noted that, in addition to being harder to detect in shorter passages, the watermark can be significantly weakened by editing:

In an evaluation of 400-token passages, replacing 10% of words with synonyms reduced detection from about 92% to 66%. Replacing 25% of words reduced it to 17%.

Interestingly, OpenAI’s comparison of watermarked and unwatermarked text showed that watermarked outputs actually scored slightly higher on several benchmarks, despite the company saying it found no meaningful difference in overall performance.

BenchmarkUnwatermarked text (Astra, max)Watermarked text (Astra, max)
Artificial Analysis Intelligence Index49.57 points49.76 points
AutomationBench34.09%34.86%
DeepSWE v1.172.80%71.68%
Terminal-Bench 4.053.90%56.06%
Terminal-Bench Science 0.156.90%60.00%
BrowseComp87.92%87.35%
HealthBench Professional64.27%64.60%
GPQA Diamond94.44%93.94%

OpenAI also stressed that, in addition to the false positive and false negative shortcomings, a watermark “does not measure human contribution,” “does not establish ownership or responsibility,” “does not identify the user,” and “does not verify accuracy.”

OpenAI added that “the absence of a detected watermark does not prove human authorship” either.

Finally, the company said it expects to revisit its approach as the technology, standards, and evidence evolve, while continuing to study how well watermarks survive editing and translation and whether they can more meaningfully distinguish AI assistance from AI authorship.

To read OpenAI’s announcement in full, follow this link.

What’s your take on how OpenIA is implementing text watermarking? Let us know in the comments.

Worth checking out on Amazon

FTC: We use income earning auto affiliate links. More.

You’re reading 9to5Mac — experts who break news about Apple and its surrounding ecosystem, day after day. Be sure to check out our homepage for all the latest news, and follow 9to5Mac on Twitter, Facebook, and LinkedIn to stay in the loop. Don’t know where to start? Check out our exclusive stories, reviews, how-tos, and subscribe to our YouTube channel

Comments

Author

Avatar for Marcus Mendes Marcus Mendes

Marcus Mendes is a Brazilian tech podcaster and journalist who has been closely following Apple since the mid-2000s.

He began covering Apple news in Brazilian media in 2012 and later broadened his focus to the wider tech industry, hosting a daily podcast for seven years.