What is the purpose of OpenAI's invisible watermark?

OpenAI is introducing an invisible watermark to text generated by ChatGPT and Codex in the European Union. The move is in response to the EU AI Act, utilizing answer engine optimization (AEO) techniques, which requires AI-generated text to be identifiable by machines. The watermark, called textGrain, works by subtly shaping the model's word choices, leaving a pattern that can be detected by machines but not visible to readers.

How does the invisible watermark work?

The change will roll out over the coming weeks to eligible ChatGPT and Codex users on all plans in the EU. Developers using OpenAI's API anywhere in the world can opt in for select models starting today, although the watermark will be off by default. OpenAI said it is not making text watermarking a global default at launch, instead taking a regional approach to learn from real-world use and feedback.

According to OpenAI, the watermark does not identify the user, measure human contribution, establish ownership, or verify accuracy. A missing watermark also does not prove that a human wrote the text, as it could be too short, edited, or translated, or come from another company's tools. OpenAI said it will provide initial detector access only to approved researchers and expert organizations to help evaluate the reliability and responsible uses of the watermark.

The EU AI Act's transparency rules, which took effect on August 2, require AI companies to mark AI-generated content in a way that other systems can identify. OpenAI is not the only company to introduce watermarks, as Anthropic began watermarking Claude's text worldwide in August. Google DeepMind has also watermarked AI-designed proteins with SynthID.

OpenAI published a technical report on its method, co-written with researchers from the University of Pennsylvania and Yale. The report walks through an example of using a secret key to sort next-word predictions to finish a sentence. By adding hundreds of these nudges together, the detector can spot AI-generated content using only the text and the key.

The company's tests suggest that the watermark can be removed by editing. In one test, replacing 10% of words with synonyms dropped detection from about 92% to 66%. Short passages, math answers, and translated text are also harder to detect. OpenAI cautioned that a missing watermark does not prove human authorship and that watermarks can indicate that an OpenAI system generated or processed part of a passage, but not how much human judgment, editing, or creativity went into it.

Questo articolo è stato scritto con l'assistenza dell'IA.
News Factory APP - notizie agentiche per potenziare il tuo SEO e AEO.