November 30, 2022 is a date that marked a before and after in the history ofartificial intelligence. It is the day when Openai launched Chatgptofficially starting a new era ofGenerative ia. Since then, nothing has been more as before. As happened on July 16, 1945, when the first atomic bomb in the New Mexico desert exploded in the United States, with irreversible consequences for the environment, also the debut of Chatgpt, according to many scholars, has Permanently “polluted” the world of data.
The analogy is strong, but not random. After the Trinity nuclear test, the atmosphere was invaded by radioactive particles that deposited everywhere, even entering industrial materials. From that moment on, no metal produced was purer, and to make high sensitivity medical or scientific tools it was necessary to resort toLow radioactive bottom steelor metal produced before 1945.
Now, in the world of artificial intelligence, something similar is happening.
Thus artificial intelligence risks self -destruction
Today, every time a generative IA produces a content – whether it is a text, an image or a code – is leaving An artificial trace in the digital environment. Traces that end in other datasets and which are then used to train new generations of models. In doing so, however, the models no longer learn from humans, but from other models. It is as if an ecosystem began to feed on only his own waste.
This phenomenon has a name: collapse of the modelor Model Autophagy Disorder (Mad). A technical term to describe a concrete risk: that the ia stops being reliablebecause its models are based on increasingly altered, inaccurate or false information.
Already in 2023, John Graham-Cumming – ex Cloudflare CTO – he perceived this danger and created Lowbackgroundsteel.aia virtual archive that collects datasets generated before the “contamination point” of 2022, such as the Arctic Code Vault, a frozen copy of the public content on Github dating back to February 2020.
Graham-Cumming’s idea? That It serves a “non -contaminated” data reservelike the steel of the past, to train future models on clean bases.
The risk of remaining without clean data
The problem, however, is wider. It does not only concern the reliability of the models, but also the equity of the system. Who still owns Human data, original and uncontaminatedcould soon have a huge competitive advantage. The startups and the small actors in the sector, however, would be forced to use polluted datasets, building models more fragile, less accurate and less sustainable.
This is the fear expressed by a group of scholars of various European universities-including the University of Cambridge, the University of Düsseldorf and the Ludwig-Maximilians of Monaco-in their paper “Legal Aspects of Access To Human-Genreda Data and Other Essential Inputs for Ai Training”published in December 2024. According to these experts, it is necessary to guarantee Public access to clean dataotherwise the artificial intelligence of the future will be in the hands of a few dominant actors.
Maurice nailresearcher at Cambridge and co -author of the study, explained the urgency perfectly:
If today we still have real human data, it is because there was a moment, as in 1919 with the sinking of the German fleet, which allowed us to keep pure steel. The same goes for data: everything that has been created before 2022 is still considered safe. But if we also lose those, we will no longer be able to go back.
We need a global policy to label and protect the original data
But how can we defend human data from the contamination of artificial intelligence? Landing the contents generated by IA is a possible solution, but. The labels can be removed, the deleted digital watermark, e Jurisdictions vary from country to country. As Chiodo recalled, Anyone can load any content on the networkand those data will then be collected and used by other models. Without control.
In their study, the authors also propose to encourage the Federated learninga system in which The data are not shared directlybut remain protected, still allowing the training of the models. A way to guarantee privacy and security, avoiding at the same time information monopolies.
However, this solution also involves risks. Who holds these data? How are they managed? What if a government that appears reliable today, will become authoritarian tomorrow?
RUPPRECHT PODSZUNexpert in competition law and co -author of the firm, underlines the importance of a decentralized and competitive management pristine data, to avoid concentrations and political influences.
Because the point is precisely this: The collapse of the models is not just a technical problembut it concerns the very future of artificial intelligence, as Chiodo warns:
If we want the IA to remain a useful, right and democratic tool, we must worry now. Because once contaminated the whole data set, cleaning it will be practically impossible.