ChatGPT will also leave an invisible trace in the texts: here comes textGrain, the system that can recognize them

In the coming weeks, a fairly long response generated by ChatGPT in the European Union may carry with it an unseen signature. No hidden characters to find with some trick, no mysterious space glued between two words: the signal will be inserted directly into the linguistic choices made by the model.

OpenAI called it textGrain and announced that it will be phased in to ChatGPT and Codex eligible text outputs in the EU. From the same day, API customers can already activate it, on some models, even outside Europe; in that case it remains optional and deactivated by default.

The objective is to adapt to the new rules on transparency ofAI Actapplicable from 2 August 2026, which require suppliers of generative systems to make content produced or manipulated by artificial intelligence also recognizable through machine-readable markings. However, the European Commission provides exceptions and distinguishes with some care between entirely generated text, simple editorial assistance and content subjected to human control.

The watermark is in the words, not under the words

The mechanism is more subtle than the word watermark suggests. When generating a text, a language model looks at several plausible words or word portions and assigns a probability to each. textGrain intervenes on this choice by introducing a statistical scheme linked to a secret key.

In the technical report published alongside the announcement, the researchers describe a system that modifies token sampling while keeping the loss of text variability under control. A detector equipped with the same key can then check whether those linguistic choices follow the expected pattern more often than they would by chance.

In short, for the reader, nothing visible changes. Copying the text into Word, pasting it into an email or removing the formatting does not cause any branding to appear because . OpenAI specifies that textGrain does not add invisible spaces, anomalous punctuation or tokens created specifically for the watermark.

The more the lyrics are rewritten, the weaker the track becomes

Here the matter becomes less convenient for those who imagine a sort of universal metal detector for homework written with ChatGPT. In tests published by OpenAI, setting the detector to have a target false positive rate of 1%, the watermark was recognized in approximately 80% of the texts of 200 tokens and in 95% of those from 400 tokens for content such as psychological answers. In mathematical texts, where the model has much less freedom in choosing words, the results are worse. They are also assessments carried out under controlled conditions, and OpenAI explicitly warns that they do not automatically describe what will happen in every real-world situation.

Then just start working on the text. In a test on passages of 400 tokens, replacing with synonyms the 10% of the wordstracking dropped from about 92% to 66%. With the 25% of words replacedfell to 17%. Deeper translations and rewrites can further weaken the signal.

There is also the language. OpenAI has verified that performance varies across the 24 official European Union languages ​​and is adjusting the intensity of the watermark precisely because some languages ​​give the model more room to maneuver than others. So that 95% should not be taken, transferred to any Italian text and transformed into a sentence.

If textGrain finds something, it doesn’t know who wrote the text

Even a positive result says much less than it might seem.

According to OpenAI, the watermark can indicate that one of its systems generated or processed at least part of the text. It cannot determine who used ChatGPT, what account did so, what prompt was written, or what conversation produced those words. It doesn’t even measure how much human work came later: a text generated and then rewritten, corrected and verified by a person can continue to contain part of the signal.

It does not even establish who the author is in a legal sense, who owns the text or whether what is written is true. And the opposite reasoning also applies: no watermark detected does not automatically mean human text. It may be too short, edited, translated, produced before textGrain was introduced, generated with an uncovered template, or come from another company.

For this reason, the text detector, at least initially, will not be put freely online. OpenAI is accepting requests from researchers and specialized organizations and will evaluate access on a case-by-case basis. The risk is quite clear: a false positive used as definitive evidence to accuse a student, an employee or a perpetrator would have much more concrete consequences than a wrong percentage on a graph.

The AI ​​Act does not impose a label on every text touched by ChatGPT

Then there is an important distinction, especially for those who use artificial intelligence in editorial work.

The AI ​​Act requires supplierstherefore, companies that develop the systems to prepare machine-readable markings for synthetic contents. For those who use these tools professionally and publish texts on issues of public interest, the obligation to clearly show that the content has been generated or manipulated by AI concerns texts that, according to the indications of the European Commission.

A substantial review carried out by a competent person, with verification of the information and editorial responsibility for the content, can therefore fall within the exemption provided. A simple automatic correction of grammar and typos, however, is not enough to magically transform into editorial control.

The Commission also specifies that the technical marking obligation does not apply when the system performs only one normal editing assistive function. Even very short fragments and code have specific rules, because there a statistical watermark has too little space to work decently.