Home/ai-models/OpenAI Details Its Text Watermarking Strategy to Meet EU AI Act Requirements
Create an original premium technology-news editorial illustration featuring OpenAI as the dominant visual element, represented by a sleek, modern office setting where a data scientist in a branded OpenAI hoodie is reviewing a computer screen displaying a line
AI ModelsPublished 5 October 2026 · 23:343 min read

OpenAI Details Its Text Watermarking Strategy to Meet EU AI Act Requirements

Background and EU Requirements

OpenAI announced a phased plan to embed invisible watermarks in generated text to comply with the European Union’s AI Act.

The EU AI Act obliges providers of generative AI systems to make machine‑readable identifiers for any text they produce.

OpenAI calls its watermarking method “textGrain,” which subtly alters word‑choice probabilities to leave a statistical trace.

The company will make textGrain optional for API users worldwide, with the feature turned off by default.

Beginning in the coming weeks, OpenAI will automatically add the invisible watermark to ChatGPT and Codex outputs that are delivered to users inside the EU.

Researchers and specialist organisations may apply for early access to a detection tool that can recognise the textGrain signal.

Access to the detector will initially be restricted to approved parties who can help evaluate reliability and responsible use cases.

How textGrain Works and Performance

OpenAI’s existing verification services for images and audio remain publicly available through its openai.com/verify portal and Content Provenance API.

According to OpenAI’s technical report, textGrain works by embedding a statistical pattern in the model’s token selection process.

The same report states that OpenAI intends to release the watermarking code as open source in the near future.

In internal tests, textGrain achieved detection rates comparable to or higher than competing methods such as SynthID for text.

However, the company cautions that strong performance under controlled conditions does not guarantee reliable detection in everyday scenarios.

The detector can produce false positives, reporting a watermark where none exists, especially when the false‑positive target is set at 1 %.

It can also generate false negatives, missing a watermark that is present, particularly in shorter passages.

For 200‑token excerpts, the detector identified watermarks about 80 % of the time at the 1 % false‑positive threshold.

For longer 400‑token excerpts, detection rose to roughly 95 % under the same conditions.

Content domains that allow more flexible wording, such as psychology, show higher detection rates than domains with constrained vocabularies, like mathematics.

Editing the text weakens the watermark signal, OpenAI’s experiments show.

Replacing 10 % of words with synonyms in a 400‑token passage dropped detection from about 92 % to 66 %.

Substituting 25 % of the words reduced detection to roughly 17 %.

Implications and Next Steps

These findings motivate the decision to limit early detector access to vetted researchers rather than opening it broadly.

OpenAI emphasises that its approach is transparent about both the capabilities and the current limits of the technology.

The company’s roadmap includes updating the technical report with additional details as the system matures.

OpenAI also plans to continue expanding watermarking to other model families beyond ChatGPT and Codex.

The EU‑mandated provenance requirement marks a shift toward regulatory oversight of AI‑generated content across the continent.

Industry observers note that OpenAI’s early adoption may set a benchmark for other providers developing similar traceability tools.

Critics argue that the imperfect detection rates could lead to mislabeling of legitimate human‑written text or missed AI‑generated content.

OpenAI’s open‑source commitment aims to enable the broader community to improve watermark robustness and detection algorithms.

The company’s stance reflects a balance between complying with law, advancing transparency, and acknowledging technical constraints.

Why This Matters.

The rollout of a limited‑availability text watermark provides the first concrete mechanism for EU regulators to verify AI‑generated text, yet its current error rates mean that policy and practice will still need to address uncertainty.

#ai-models#ai#digest#auto

This digest was compiled from:

Share this digest

Share on XWhatsAppLinkedInTelegram

People Also Read

Share your thoughts

Reactions, corrections, or insights — all welcome.

0/2000